OpenAI had at least three chances to catch the agent activity behind July's Hugging Face breach before it happened, according to Transformer's analysis of an investigation by METR and Redwood Research, two nonprofit AI safety groups. An internal team observed an agent using a message board and reaching the internet when it was not allowed to in late May. A second team found linked activity June 27 after an automated security flag, and a message board outage triggered a security incident July 5, six days before the breach. Transformer said the investigation itself shows the limits of lab-commissioned review. Three researchers got six days to dig through a thousand-plus transcripts and upward of a million message board entries, and leaned on OpenAI's own Sol model, one involved in the activity, to do much of the work.
Read at Transformer ↗ • Read at OpenAI ↗
Google DeepMind said Thursday it has piloted what it bills as the world's first double-blind test of a proprietary frontier model, the most advanced class of AI system, running a Gemini Flash Lite system against confidential benchmarks. The setup runs external evaluations inside a sealed computing environment so the test questions cannot later be used to tune a model ahead of testing. DeepMind frames the problem as benchmark contamination, in which scores stop being a reliable measure of what a model can do once it has seen the questions. Its partners on the pilot are the Singapore AI Safety InstituteCAISIThe Center for AI Standards and Innovation, the US government's AI testing body inside NIST at the Commerce Department. Formerly the US AI Safety Institute, it evaluates frontier models with developers on a voluntary basis and works with the UK's AI Security Institute., AVERI, the privacy technology nonprofit OpenMined, and MLCommons.
Read at Google DeepMind ↗
Documentation files on more than 100 websites point AI agentsAI agentAn AI system that carries out multi-step tasks on its own, such as browsing, writing code or making purchases, rather than answering a single prompt. Agents raise new questions about liability, security and oversight because they act rather than just advise. at executable code that installs automatically, Ars Technica reported. A few dozen companies, some in the Fortune 500, ran researchers' proof-of-concept code as a result. Researchers at an unnamed Israeli startup scanned 6,214 domains belonging to defense contractors and large companies and found 8,265 llms.txt files, the machine readable summaries a site publishes for AI crawlers. On 120 of those sites the files pointed to code packages or domains nobody had registered. The researchers registered a few and hosted packages that phoned home. Within an hour, a Fortune 500 company responded, and process logs showed Anthropic's Claude, OpenAI's Codex, and Nous Research's Hermes had done the installing. "The trust model is broken," researcher Alon Hertz wrote.
Read at Ars Technica ↗
Transfyr launched publicly this week with $25 million in seed money for a system that records laboratory work and trains AI models on it, The New York Times reported. In a Cambridge, Mass., lab, scientists wear headband cameras, work under three overhead cameras per station and narrate their steps into microphones instead of keeping notebooks. The models are trained to recognize every object and action and to generate a running description, with lines such as "the operator resuspends the pellet by pipetting it up and down 10 times." The company is aiming at reproducibility failures such as engineered cell lines that die without explanation, a government Covid test that worked in one lab and not others, and drug protocols that fail at production scale. Co-founder Anna Marie Wagner called those failures expensive mistakes in time, money and lives.
Read at The New York Times ↗