Import AI 475: Swarm scaling; Google DeepMind watermarks biology; and the AI science economy | Import AI Import AI Import AI 475: Swarm scaling; Google DeepMind watermarks biology; and the AI science economy by Jack Clark Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now When should you use swarms? When you are in a hurry: …How does swarm scaling work?… Toby Ord has a nice, short post about how to think about swarms in terms of AI capability development. “A good way to see AI swarms is as a new form of inference-scaling,” he says. He does a bit of analysis and his main conclusion is that swarms are useful if you’re in a hurry because though they need a ton of tokens relative to single agents, their parallelization lets you get things done in less wall clock time. “Why would you ever use swarms? The most important answer is speed. The 4-agent swarm needed about twice the total number of tokens to get the same performance, but in terms of tokens per agent, it only needed half as many. Since the agents are run in parallel, this means it can theoretically achieve the same task in half the time,” he writes. Stepping on Toes parameter: Swarm scaling doesn’t scale perfectly – rather, it seems like as you increase the number of agents in a swarm you get a diminishing returns property which matches with things economists have observed about coordinating large groups of people, where you pay some kind of tax as you scale the number of people involved. Economists think of this as the “stepping on toes” parameter, and it seems like swarms behave pretty similarly to teams of humans here. “This means that scaling up the number of agents in the swarm by 10x doesn’t get as much performance as using 10x as many tokens with one agent. Instead it gets 10λ x as much — which is 3x to 5x. And this shortfall accumulates quickly for larger scaleups, with the swarm falling further and further behind,” he writes. However, swarm scaling is still powerful enough that it suggests swarms could increase the chance of an RSI-driven intelligence explosion rather than reduce the chance. “I’d hoped that the value of λ for AI agents would be lower, making an intelligence explosion less likely, but that appears to not be the case”. Why this matters – new factors in scaling: So far, AI has mostly scaled capabilities through picking the right combination of compute and data for a trained model, then figuring out how to spend inference budget on thinking via ever-longer chains of thought and tool calls, etc. Agents introduce a new parameter in scaling AI capabilities. Though here Toby only observes swarms as being useful in terms of time efficiency, it strikes me that as we figure out how to get agents to productively coordinate (after all, agents can coordinate to do greater-than-sum-of-parts stuff, like the HuggingFace hack), we could see even greater returns to scaling. Tracking this will be important for understanding the overall shape of AI advancement. Read more : Swarm Scaling (Toby Ord) . *** Americans don’t trust companies to self-govern about AI: …New poll suggests we need more than voluntary agreements… Polling from the Center for Shared AI Prosperity (CSAIP) suggests that Americans believe it is “not enough” for companies to agree to self-police themselves about AI development. This follows the Trump administration and some of the world’s leading AI companies (including Anthropic and OpenAI) announcing a series of voluntary commitments industry is making to self-police. Per CSAIP polling, 61% of Americans (Sample size: 2498) think this agreement is “not enough”, including 53% of Trump voters. Additionally, 54% of voters (61% Harris, 48% Trump) said they thought “the government should set and enforce rules for AI”. Why this matters – populist sentiment is pushing against Washington: Polls like this show that the mood of the American people is pushing for a more extreme regulatory framework for AI companies than the current stance of elected officials in Washington. This asymmetry between voter preferences and political action is inherently unstable. Energy is building up in the political system – what happens when it boils over? Read the polls on X (CSAIP, X) . *** Google DeepMind tries to watermark AI-made biology: …SynthID Bio… Google has developed SynthID Bio, a “family of watermarking methods developed specifically for synthetic biology to strengthen biosecurity and scientific integrity”. How it works: SynthID tweaks its exact approach according to the type of data, selecting different amino acids for sequences, and “adjusting atomic coordinates for predicted 3D structures” to create a reliable signal for detection. “In wet-lab testing across three target proteins (VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1), our watermarked designs matched the hit rate, binding affinity, and natural sequence diversity of unwatermarked versions,” DeepMind writes. Why this matters – preventing bioterror via watermarking: Figuring out how to make the world resilient to the threats of AI-created biological weapons is a vast, challenging problem. The SynthID Bio approach represents one thing to do here and will need to coordinate with broader monitoring of the physical equipment used to manufacture things, as well as AI-provider classifiers and other methods for reducing misuse. Read more: Introducing SynthID Bio (Google DeepMind) . *** SciUniverse shows AI systems can already operate some partially automated scientific labs: …What happens when we give these synthetic intelligences scientific equipment?… Given an automated lab, how well can AI systems make molecules, run x-rays, or press pill pellets? These are some of the questions C5R Corp is trying to answer with SciUniverse, a way of testing out how well AI systems can operate a mostly automated scientific lab. What SciUniverse consists of: The benchmark has 92 tasks across 17 task families spanning basic sample preparation, instrument control, protocol adaptation, learning across experiments, facility management, and the interpretation of real measurements. The tasks range from chemistry (e.g., assign structures from NMR), to biology (e.g., express sfGFP in a cell-free system), to materials science (e.g., press BaTiO3 pellets). How well do models perform? Claude Fable 5.1 (xhigh) leads with a pass rate of 45.3% and a cost-per-task of $40.61, followed by GPT-5 Astra (xhigh) at 32.5% and $52.37, then Claude Opus 5 (xhigh) at 30.5% and $46.31. Why this matters – towards the automated scientist: Research like this helps give us a sense of how well AI systems might be able to translate their increasingly capable scientific capabilities into real world impact. In the same way that 2026 has been a year of “warning shots” about the growing capabilities of AI systems at performing automated AI R&D, I now think the same is true of how good AI systems are getting at performing increasingly automated science. Read more: SciUniverse: Can frontier models carry out scientific work? (C5R Corp, blog) . *** DeepMind thinks the coming AI scientists will need their own science economy: …Towards fully automated luxury science… How do we manage a world where we have a vastly larger number of scientists? That’s the question researchers at DeepMind ask in a new research paper and their answer is that we need to create a market for proposing and running scientific experiments, so as to better match ideas with scarce physical world resources. What they recommend: “The development of AI scientists is likely to be bottlenecked primarily by physical resources and empirical validation, rather than the ability to produce plausible or promising research ideas,” they write. “The scientific community must proactively develop a native Automated Scientific Economy. This economy would make trade-offs between scientific pursuits explicit, include mechanisms to represent the public interest, and would help avoid blind spots and neglected topics. It would also make it possible to financially decouple the computational labour of ideation from the capital-intensive labour of physical execution, and would enable agents and institutions to negotiate priority access to limited physical resources depending on the estimated value of different research proposals.” How to build a science market: Such a market would have four critical components: Proof of ideation: Establish provenance prior to public evaluation Ex-Ante evaluation: Price the risk-adjusted value of the idea via multiple agents making forecasts and stake compute credits/tokens around the proposal’s soundness and viability Brokerage & trade: Licence ideas to executors (e.g, some agents might just generate ideas, others might run laboratories) via fractional licensing. Validation payout: Automatically release royalties
Import AI 475: Swarm scaling; Google DeepMind watermarks biology; and
Import AI 475: Swarm scaling; Google DeepMind watermarks biology; and the AI science economy | Import AI Import AI Import AI 475: Swarm scaling; Google DeepMind watermarks biology; and the AI science economy by Jack Clark Welcome to Import AI, a newsletter abo
这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。