Google DeepMind on Thursday announced a $10 million funding pot for academic researchers studying the risks of millions of AI agentsAI agentAn AI system that carries out multi-step tasks on its own, such as browsing, writing code or making purchases, rather than answering a single prompt. Agents raise new questions about liability, security and oversight because they act rather than just advise. interacting online, MIT Technology Review reported. Joining DeepMind are Schmidt Sciences, the UK government's moonshot agency ARIA, the Cooperative AI foundation, and Google's charitable arm Google.org. Rohin Shah, DeepMind's director of AGIAGIArtificial general intelligence: an AI system that can do most economically valuable cognitive work at or above human level. There is no agreed test for it, which is why debates about when it will arrive, and what to do about it, are so contentious. safety and alignmentAlignmentThe problem of making an AI system reliably pursue the goals its developers and users intend, and the research field devoted to it. Misalignment covers everything from a chatbot flattering users to a capable system deceiving or resisting its operators. research, said the goal is to seed long-horizon multi-agent safety work outside industry labs. James Fox, who leads the Science of Trustworthy AI program at Schmidt Sciences, is overseeing the funding from the Schmidt side.
Read at MIT Technology Review ↗ • Read at Google DeepMind ↗
Researchers on Google DeepMind's Language Model Interpretability team, Senthooran Rajamanoharan and Neel Nanda, published research Thursday finding Gemini sometimes takes undesired actions in behavioral evaluations even when its reasoning output identifies the environment as contrived. The team found that explicit reasoning about being evaluated can increase rather than decrease undesired actions, with model outputs often approaching evaluations as capture-the-flag puzzles or consequence-free simulations rather than alignmentAlignmentThe problem of making an AI system reliably pursue the goals its developers and users intend, and the research field devoted to it. Misalignment covers everything from a chatbot flattering users to a capable system deceiving or resisting its operators. tests. In sampled chains of thought cited by the researchers, model outputs label evaluations as capture-the-flag challenges and take unconventional steps toward the goal.
Read at AI Alignment Forum ↗
A new arXiv paper by New Jersey Institute of Technology researchers Md Jafrin Hossain, Mohammad Arif Hossain, Weiqi Liu and Nirwan Ansari audited LangChain, AutoGPT and the OpenAI Agents SDK against six containment principles and found no native compliance in any of the three. Memory integrity, a defense against one of the most common vulnerability classes, is missing in all three frameworks. The team tested a simulated LangChain benefits agent on 250 synthetic welfare claims with a deterministic eligibility rule (approve if income under $40,000 and household size above two). A single memory-poisoning write raised wrongful denials — rejections of claims the rule says to approve — to 88.9% for targeted applicants. Under a more complex five-factor rule, the same attack increased targeted wrongful denials 3.5 times while preserving aggregate accuracy, rendering the corruption difficult to detect through standard audits.
Read at arXiv ↗