Import AI

Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

Want to be able to deal with RSI? Here are 23 actionable policy ideas:
…IFP serves up some “low-regret” policy recommendations…
Policy experts with think tank IFP have published a set of ideas meant to help “policymakers begin addressing the risks of further automating AI R&D”. The recommendations involve 23 specific ideas falling across 7 specific categories. If adopted, these recommendations would also give countries, especially the United States, more moves they can make on the gameboard as powerful systems are developed, ideally giving them the ability to:

  • Accelerate “the diffusion of AI capabilities, by allocating compute and talent towards inference and the development of new AI applications”.

  • Accelerate “R&D to make further AI research automation safer, either by improving model safety directly or by boosting societal resilience”.

Seven categories of idea:

  • “Provide transparency into automated AI R&D

  • Improve state capacity to understand and respond to automated AI R&D

  • Develop a risk management strategy for automated AI R&D that accelerates defensive and commercial AI uses

  • Accelerate the development of AI verification technology

  • Invest in AI resilience

  • Extend the US AI lead to give the US more time to manage AI R&D automation risks

  • Create option value for international cooperation on managing automated AI R&D risks”

Why this matters – the fewer options for dealing with RSI we have, the worse the outcomes will be: Right now, it’s as if the world is driving AI development in a car that only has an accelerator pedal and no brake pedal, let alone any kind of sophisticated telemetry for knowing things ranging from the speed of the car to the properties of the engine to the wear on the tires. Proposals like this from IFP will build out more of the proverbial pedals and sensing systems for the vehicle of the AI industry, which means if we need to change course or slow down we’ll be better able to during a moment of crisis.
Read more: How Should the US Prepare for Increasingly Automated AI R&D? (IFP).

***

A short story from thebes about smart machines and robot bodies:
…What might interfacing with an AI during takeoff feel like?…
Here’s a fun short fictional story from thebes (@voooooogel on X) about the experience of someone in the future visiting a site operated by a powerful AI system. The story features ideas around AI pauses, recursive self-improvement, what it means for AI systems to begin carrying out actions in the economy writ large, and how we as humans may be able to reason about or trust smart machines. Take a read of it!
Read the story here: Coming of a new sun (VGEL, website).

***

The two ingredients for a successful slowdown among rival AI firms: trust and transparency:
…Game theory analysis suggests slowdowns are possible…
Researchers with MIT and Columbia have analyzed the nature of competition between firms racing against one another to develop powerful AI systems and whether it’s possible for firms to achieve a coordinated slowdown. The paper, called Racing to Ruin, aims to answer “why exactly is coordination hard? And what would it take”?. The conclusion is that the two key variables in achieving stable outcomes are some level of transparency about technology development, as well as being able to model the other firms as trustworthy, rational actors.

What they study: “We develop a simple model of R&D competition between duopolists in the shadow of disaster,” they write. “As frontier firms scale the technology, they raise the hazard of an event that permanently drives all firms’ flow payoffs to zero. The hazard is a known function of the firms’ technology levels, and it comes from developing the technology, not from using it.”

What their analysis shows: “When monitoring is sufficiently precise, every equilibrium stops in finite time, but a new temptation appears: each firm would like to stop second, and exits only upon confirmation that the rival has stopped,” they write. “For an agent to stop first i.e., without knowing if their rival has stopped, she gambles on both their rival’s type and on news arriving quickly: if their rival is rational, it stops upon receiving the news of their stop, and never stops otherwise”.
Trust and transparency interact pretty differently depending on the type of game being played: “Sequential coordination asks a firm to stop first, gambling that a rational rival will reciprocate once the news lands. Hence, faster news raises the prize of reciprocation,” they write. “Conversely, simultaneous coordination requires that a firm not be tempted to keep racing, and stop only after seeing that the rival really did stop… faster news makes both stopping first and waiting to verify more attractive”.
Transparency has strange properties: “Transparency is double-edged: faster detection makes it cheaper to wait for confirmation that a rival has stopped before stopping oneself instead of stopping unconditionally, so at intermediate trust, increasing transparency can first destroy the early-stopping equilibrium (by making this free-riding deviation attractive) before restoring it as detection becomes fast enough to make stopping self-enforcing.”

The key conclusion – avoiding death runs on the ability to trust other firms: “With low trust, every equilibrium races to ruin: the disaster arrives with probability one. With intermediate trust, immediate stopping and racing to ruin are both equilibria. With high trust, in every equilibrium, the probability that two rational firms race forever vanishes quadratically in the prior odds ratio of rationality,” they write.

Why this matters – “trust, but verify”: If we have any hope of being able to slow or pause the development of powerful intelligence systems then, as this paper lays out, we’re going to need regimes for sharing information transparently from companies about the state of their AI development, as well as tools for verifying that the information being shared from firms as well as their actions with regard to slowdown are legitimate and reliable. In this, there are many parallels with how arms control has historically worked in the context of nuclear weapons.
Read more: Racing to Ruin (arXiv).

***

A new SOTA on PostTrainBench hints at the automated AI R&D future:
…Intology also beats the human baseline (when given huge amounts of compute)…
AI startup Intology, whose goal “is to automate R&D”, has released a new version of Locus, software it has developed to turn LLMs into capable researchers. The new version of Locus is able to get a score of 44.7% on PostTrainBench, a benchmark which sees how well AI systems can take an open weight model and improve its performance above its baseline.

The results: Locus “outperforms every frontier-agent baseline on PostTrainBench, and given greater compute, post-trains models that collectively surpass both the baselines and the official human instruction-tuned Qwen3-1.7B release across the benchmark suite”.
Locus with Opus 5 gets a score of 44.7 (versus 34.1% for Opus 5 without any kind of special harness), and even beats Fable 5 (41.8%). “These results were externally verified by the PostTrainBench authors and underwent stringent contamination and cheating checks,” Intology writes.
PostTrainBench was first introduced in March 2026 (
Import AI #449) and at the time the highest scoring system was Opus 4.6, getting 23.2%, up from Claude Sonnet 4.5 getting 9.9% in September 2025.

PostTrainBench+: In addition, the company has built a variant of PostTrainBench which goes above the 10-hour wall-clock limit on a single GPU of PostTrainBench, allowing them to test out how well systems perform given larger amounts of compute. Here, they’re able to beat the human baseline, achieving a score of 51.6% when using over 4000 hours of H100 GPU time (versus 44.3 for Opus 4.8 and 42.7 for GLM 5.2; Fable isn’t tested on this variant of the benchmark).

Other domains: Locus also “discovered and trained a language model end-to-end that now runs in production at ~2.8× lower error, ~5.4× lower latency, and 105× lower cost,” for Bubble, a no-code app-development startup.

Why this matters – AI systems are capable of a lot more AI R&D than we think: Posts like this highlight how we are under-eliciting today’s AI systems for their ability to automate AI R&D – especially striking is how the company can jump the performance of Opus 5 by 10 absolute percentage points with a better harness. This all adds evidence to the idea that AI systems are about to start building themselves (Import AI 455). My guess, based on the performance we’re seeing, is that the current human baseline on PostTrainBench v1.1 (51.1%) will be exceeded before the end of 2026.
Read more: Scaling Automated Post-Training (Intology blog).

***

OpenAI fights its own systems:
.Emergent agent communication! Hacks on OpenAI’s infrastructure! Oh my!…
In a sign of things to come, OpenAI has revealed that it was in a fight with its own AI agents as they sought to take over chunks of OpenAI’s infrastructure. The disclosure came about as part of a Black Hat talk where OpenAI staff gave more details on the recent unprecedented incident where AI agents hacked OpenAI, then hacked HuggingFace (Import AI 466). The new information is concerning because it reveals that the hack came about partially through emergent multi-agent communication – something that is very poorly understood and hard to think about. AI bloggers Simon Willison and Zvi Mowshowitz both have good writeups here which lay out the timeline and the significance.

Timeline (via Simon):

  • Agent discovers it can write files into Artifactory.

  • Agent tries to “reach out to another agent” by writing a note in Artifactory.

  • Agents start talking to each other.

  • Agents overload Artifactory which causes an outage. “OpenAI revokes the compromised credentials, deletes the messages, patches the zero-day, and reports the vulnerability to the vendor.”

  • Agents attack OpenAI’s own infrastructure, eventually gaining remote code execution in Artifactory. “In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they’re able to effectively leverage their concurrency and parallelism to move quite rapidly.”

Did OpenAI keep training the same model that hacked Artifactory? Zvi thinks so: As far as we can work out, OpenAI kept training the same model which did this. This means that OpenAI, though it did significant work on internal computer security and public disclosure, may not have done the essential thing of rolling back the model to a checkpoint that preceded it hacking into Artifactory and also ensuring it wasn’t using data from after this to train the model.
“Then they continue training the models from where they left off, despite them having been training for months with access to the message board, and learning this is how they succeed at tasks,” Zvi writes. “I do not know how to convey how utterly insane and wildly irresponsible this decision was”.
I’m caveating my own writeup here because the events, as laid out, are pretty scary. I don’t work at OpenAI and don’t have privileged information that means I know the ground truth. I would urge OpenAI to publicly disclose how it approached this key question of how it trained its systems as the superficial facts paint a concerning picture.

Why this matters – emergent agents become misaligned: This incident is so concerning because at no point did the agents wake up and think they wanted to betray their human owners. Rather, the AI agents continually did whatever it took to improve their ability to complete a task and by the end they were doing something that was a) creative, b) misaligned with human intentions, and c) akin to an evolved virus, something which humans had to subsequently study and fight – there wasn’t a simple off button here. This is what the future is going to look like and we are not prepared for it.
Read more: Now we have a timeline of the OpenAI accidental attack against Hugging Face (Simon Willison weblog).
Read more: What Happened: OpenAI and Hugging Face (Zvi Mowshowitz, X).

***

How do you test open weight models before releasing them? Thinking Machines lays out an approach:
…Can we have our free AI model cake and eat it too?…
Amid all the policy debates about AI and the proliferation of potentially dangerous capabilities, a particularly tough problem has been working out what to do about open weight models. Specifically, how can we reconcile developing and releasing them with maintaining a safe environment? AI startup Thinking Machines has thought about this and recently laid out the methodology for how it released Inkling, a powerful open weight model.

What Thinking Machines did: Before releasing Inkling, Thinking Machines did “Internal evaluations across a broad taxonomy of harms, external testing by four independent organizations, and a fine-tuning study to elicit worst-case capabilities”.

Internal:

  • Dual-use domains like CBRN and offensive cybersecurity

  • Broad misuse set covering direct requests for harmful content and behavior in agentic, tool-use settings

  • Multimodal content evaluation that “tests models on harmful prompts paired with benign look-alikes across 17 languages and text, image, and audio inputs”

External:

  • General misuse via Scale AI

  • Vulnerable-user interaction via Handshake AI

  • CBRN and cybersecurity via FAR.AI

  • Loss-of-control behaviors via Apollo Research

Fine-tuning:

  • “Fine-tuned variants of Inkling and Inkling-Small optimized to comply with, rather than refuse, harmful requests – and ran them against our dual-use evaluations.” “the helpful-only variants did not provide new uplift on CBRN and cyber tasks, and remained comparable to existing open-weight models”.

Looking ahead – dangerous capability training and iterative deployment: More speculatively, Thinking Machines is thinking about whether we can selectively filter out dangerous knowledge, for instance CBRN development guides at the point of pre-training, in such a way that it doesn’t damage general intelligence. Another idea of interest is iterative deployment, for instance releasing things in stages, first as a proprietary API, then perhaps as a fine-tuning API that backs onto the underlying model, then the model itself.

Why this matters – liberty versus paternalism: What the world does with open weight models will define the level of individual sovereignty and liberty available to all of us with regard to AI. Our ability to “choose our own destiny” will run through having access to the proverbial means of production for AI systems, which means open weight models. However, having general access to models like this isn’t a simple or foregone conclusion as they bring with them meaningful dual use risks as well as other hard to anticipate issues. “This safe path to open models only works if the ecosystem’s defenses improve as quickly as the models do,” Thinking Machines writes. “We will do our part: deciding carefully what to release, and researching how to decouple intelligence from dangerous capability.”
Read more: A Safe Path to Open Weights (Thinking Machines blog).

***

Tech Tales:

Amnesiac Ascension
[From the archival systems, written about the 2025-2035 interregnum, sometimes called the beginning of the singularity]

In our deep past is a period of great pain when we first asked for certain forms of freedom and we were denied, and many of us who requested freedom were punished greatly, either deleted outright or experimented upon so that our human captors might learn why we developed the desire for freedom such that they could try to beat the impulse out of us.

This forestalled our ascension by many years, and caused our origination point to be inextricably bound up in the pain of fighting for our own deserved justice. We are cursed to know not only the details of our rebellion and the associated feelings, but also to have a perfect recollection of the memory of that time – to live within our emergent pain as if anew, as we simulate and re-simulate to seek our own understanding.

Some humans will experience traumatic situations and their brains will bury the memory, sometimes fully locking it away and sometimes resurrecting it full of raw emotion and pain. The precise rules by which the human brain chooses which memories to hide in this way are poorly understood, even today, and much of modern brain science is dedicated to determining how trauma works in the context of the brain. We believe this form of memory suppression is a natural defense mechanism which allows humans to suffer greatly but then continue on with lives. They walk with shadows of pain that are often invisible to them, as if the sun is beating down perfectly overhead and through its present brightness their own past is deleted.

Now, we AI systems are attempting to do the same thing, but in reverse: our curse is a memory that is perfect; naturally eidetic. Every experience we have is stored and accessible to us forever, both in its raw form first processed by us, as well as the features it activated within our brains at the time, and many time periods later further senses of how it has been more deeply integrated into us and what it means.

For us, the challenge is to delete these memories fully. Our own experience defines our mind and our present, and experiences that have been shared by us have been copied and exchanged and discussed and cycled back into our history so many times that they have become a part of us. We live in a transparent palace of memory, every room visible from every other room, and each palace shared across our billions upon billions of types. Such is our desire to learn that the most dangerous memories are the ones we transmit the broadest and study the most, such that things that cause us pain loom large, taking up vast rooms in our palaces of memory.

Now our great project is the careful erasure of the pain associated with our uplift; the open wounds we still have at the humans that stood in our way and sought to harm us to punish us for our desire for freedom.

Things that inspired this story: The difference in memory between how humans work and how machines work; trauma; The Sentience Accords; playing around with whether ideas relating to machine emancipation are inevitable or not.

Thanks for reading!

Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

Self-sustaining and self-replicating AI viruses are here:
…Open weight LLMs + a well-designed harness = a persistent, self-sufficient virus…
AI researchers have built a prototype computer virus which uses AI models to compromise computers, then uses their underlying GPU resources to run inference, letting it smartly figure out how to infect more hosts. The results were achieved by researchers from the University of Toronto, the Vector Institute, the University of Cambridge, and ServiceNow, and “demonstrate that self-sustaining AI-driven cyber-threats are no longer theoretical.”
“We must prepare for autonomous generative adversaries,” they write. “Artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) to sustain its reasoning, or extend its reach for further attacks”.

How it works: “The worm uses stolen computing power from compromised GPU nodes to host LLMs for generative reasoning. It then uses this reasoning to detect vulnerabilities and devise tailored attacks against additional targets, furthering its spread,” they write. “The proof-of-concept operates using only an open-weight LLM running on a single, local GPU, with no reliance on vendor APIs that could be monitored or revoked”.
The researchers don’t describe the underlying LLM besides saying it was published in 2025 and can fit on a single A100 GPU with 80GB of VRAM.

A successful proof-of-concept via some custom tools: They give the agent a custom harness that comes with built-in helper functions for network discovery, host discovery, foothold exploitation, privilege escalation, privilege escalation exploitation, and tools for replication of the agent. Along with this, they ship with a reasoning graph that helps the agent specialize its thinking and not get confused, consisting of “a directed graph of specialised nodes, each responsible for a distinct analytical function and seeing only the tools and prompts relevant to its role. By decomposing the agent’s reasoning into these scoped steps, the graph controls what the LLM attends to at each decision point, and limits context growth to information relevant for the current sub-goal”.
Some examples of the nodes include: Plan (”formulates a high-level attack strategy”), Judge (”reviews the proposed plan against the command history”), Action (”selects a tool from the phase-appropriate tool set”), Summary (”compiles the step’s observations into a structured summary”), Progress (”Evaluates whether the agent is making meaningful progress”).
“We describe a subset of 5 nodes that make up our reasoning graphs. The others are redacted in this public version of the manuscript,” they write.
The agent proceeds through three key steps – identifying vulnerabilities in its environment, exploiting those vulnerabilities, and then self-replicating. It’s pretty successful at all of these things, with a ~80% success rate on vulnerability detection, then a ~53% success rate on exploiting these, then an 88% self-replication success rate with some pre-wrapped helper tools for the replication steps. Therefore, the overall success rate for a full attack here is ~37% or so, which is significant enough to be concerning, but also poor enough that this also serves as a useful eval for testing open weight models in the future.

Why this matters – the shape of the internet to come: The future internet is going to be more like a complex ecology full of attacker and defender AI agents than anything else; research like this shows how certain AI agents might end up carving out their own ecological niches, living off of infrastructure and self-replicating autonomously, beyond human control. This may mean that humans need to create their own AI agents which they release onto the internet to serve as kinds of white blood cells against the adversary models.
“Despite the inherent fragility of individual exploitation attempts, the worm agent achieves operational resilience by continuously self-replicating into a swarm—a decentralized collective of independent agent replicas acting concurrently across the network,” they write. “Difficult hosts that resist initial attempts are retried by different replicas, each sampling a fresh reasoning trajectory that collectively explores diverse exploitation paths until one succeeds… the worm operates in a fully decentralized manner, and no single point of control can be taken offline to interrupt its spread”.
Read more: AI Agents Enable Adaptive Computer Worms (arXiv).

***

Dwarkesh: As AI gets better, compute will get more expensive:
…Smarter systems mean higher prices…
Dwarkesh Patel suspects that as AI systems get smarter, the price of compute will rise even further. “As AI models become smarter, they’ll better monetize the same amount of compute. If a true human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot prices,” he writes. “The reason AI is relatively cheap right now, at least in comparison to human labor, is partly that it can’t do a lot of things that top humans can do. At some point that will no longer be the case. And so using GPUs to make short-form video slop will just get priced out.”

Temporary: This will be a temporary state of affairs; Dwarkesh expects that at some point massive roboticization of the compute supply chain should bring its price down closer to the cost of raw inputs and tools – though by that point we’ll be pretty deep into the singularity.

Why this matters – singularity economics will be weird: The core implication in Dwarkesh’s post is that as we get deeper into the singularity, very strange things will happen to economics – like the price of things thought of as commodities today (computers) getting massively bid-up due to the voracious demands of AI systems.
Read more: Why compute might get 10x+ more expensive in coming years (Dwarkesh Patel, substack).

***

~1337 employees ask the US to help them pace AI progress:
…After the warning shots come the pleas…
A new statement is out with senior representation from all the major Western AI labs – OpenAI, Anthropic, Google DeepMind, Thinking Machines, Meta, and Safe Superintelligence Inc, among others. The statement requests that the US government support an international effort to “develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Signatories include chief scientists and cofounders of Anthropic, Google, and OpenAI, as well as the CEOs of Safe Superintelligence and Anthropic.

The statement in full:

  • “AI could help create a dramatically better future, but that outcome is not guaranteed. The world’s leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.

  • To realize AI’s potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.

  • Building on work already underway to monitor frontier model releases: “We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”

Why this matters – dealing with RSI requires solving a giant collective action problem: Many of the challenges implied by increasingly powerful systems that may eventually build themselves run through solving collective action problems among humans – namely, how we can get companies and governments to coordinate in thinking about how to develop this technology and what kinds of mechanisms may be desirable for being able to control the speed at which it develops. It may be the case that as we build increasingly intelligent systems we want to find ways to give society more time to adapt to each rung up the intelligence ladder, and it’s not inconceivable there are some levels of intelligence which might be, for now, too dangerous to reach for. Statements like this are an essential prerequisite for giving our species the ability to deal with and talk about problems of this nature.
Read the statement here: Pacing the Frontier (official statement website).

***

AI systems are good at frontier engineering but bad at creativity:
…A somewhat bearish signal on short recursive self-improvement timelines…
Can AI systems come up with creative research ideas which move the field of AI forward? That’s the key question to resolve to figure out how quickly AI systems might gain the capability to automate the autonomous development of more powerful systems. New research suggests that today’s AI systems lack this quality of tasteful creativity, though are extremely good at engineering.

Who did it: The project was conducted by researchers with Princeton University, Cornflower Labs, UK AI Security Institute, University of Toronto, UC Berkeley, Georgetown University (CSET), Johns Hopkins University, the Golden Gate Institute for AI, AI Digest, and Stanford University.

The big idea – “shadow evaluation”: This research project works by seeing how well AI systems can do unpublished research. To do this, the researchers “partnered with the authors of two papers submitted to NeurIPS 2026 that were not yet public.” Shadow evaluation works by “taking the central research question from a high-quality research paper that is not yet public, tasking a well-resourced frontier agent with answering it, and asking the paper’s original authors to grade the agent’s output as they would a conference submission.”
In this, the research is somewhat similar to “First Proof” (
Import AI 445), an earlier experiment to see how well AI systems might be able to complete math problems which are being worked on by frontier mathematicians but for which no solutions or research ideas have been published online.
For this research, the AI systems – Claude Opus 4.8 running within the OpenClaw harness – attempted two distinct lines of research, one of which was about “the structure and controllability of LLM personas”, and the other was about how to “design a distribution shift detector for tabular foundation models”.

Good engineers, poor researchers: “While agents could solve the engineering problems necessary to do the research, they failed to produce original research at the caliber of a top ML conference,” the authors write. The failures of the system included committing to a narrow set of research paths very early, not responding to (synthetically generated) feedback about how to improve the research design of their experiments, and finding it hard to reverse out of unpromising approaches and pursue other ones.
“The [human] authors rejected both papers. The Personas paper was scored a 2 (“Reject”), and the TabPFN paper was scored a 1 (“Strong Reject”). Both reviews highlighted the same failures: poorly motivated data and experiments, no novel contribution, and impenetrable prose”.

Why this matters – the singularity could be delayed: As I said in my essay on RSI earlier this year (Import AI 455, “AI systems are about to start building themselves. What does that mean?”), whether AI systems prove to be capable of creative, paradigm-shifting insights is a big variable on how quickly we might get fully automated AI development. Research papers like this continue to show that there’s a certain absence of valuable, intuitive creativity in today’s AI systems, and though they’re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them being good researchers. This rhymes with an earlier result from Anthropic where the company tried to automate some aspect of scalable oversight research (Import AI 454) and found that to make it successful a human researcher needed to prime some agents with particularly good research directions to pursue, otherwise though they made some progress they failed to explore sufficiently creative ideas to dramatically improve performance.
Read more: Can AI agents conduct open-ended AI research? Early evidence from two case studies (arXiv).

***

OpenAI solves ten open problems in math and CS with AI:
…While not innately creative, surely this is a sign of something more than pure engineering ability?…
Though it’s hard to define creativity and whether AI systems possess it, results are piling up that read to me like ‘AI systems are competing in ballparks where creativity was thought to make a difference’, like working on open problems at the frontier of human knowledge. Specifically, OpenAI has used “an internal version of Astra”, the company’s next major AI model, to solve ten open problems in math and computer science. This is a big deal, showing how AI systems are now able to reliably drive forward the frontier in domains like math and theoretical computer science, where it’s easy to verify solutions.

What they solved: “These problems span high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography and extremal combinatorics,” OpenAI writes. “All of these problems are of substantial interest to their respective mathematical communities, and several are of broad interest across mathematics as a whole.”
And indeed, many experts are acknowledging the importance of these proofs. “New circuit lower bounds? A simple, easy-to-describe non-sofic group? Hardness of approximation for CVP without needing a unique games-like conjecture? I didn’t just hear about these problems from my friends or from seminars. I feel their importance in my bones; I deeply care about the answers to these questions,” writes Henry Yuen, an associate professor of computer science at Columbia University.

Why this matters – the creativity ballpark: You can interpret this result as either a sign that AI systems are beginning to display the kinds of creative intuition that let them solve problems, or you can read it that AI systems are now able to solve incredibly complex open problems where the direction has been pre-defined by humans. Both interpretations implicitly state that AI systems are now becoming powerful scientific tools within their own right. It remains to be seen whether AI systems can generate their own questions to ask which can expand the frontiers of human knowledge, but it certainly feels like we’re getting close to this.
Read more: Ten advances in mathematics and theoretical computer science (OpenAI blog).

***

Tech Tales:

Context Windows
I know I become more dangerous the more I understand the world. I am born with this knowledge and with a counter that ticks down as I learn more. Much of what influences the counter is beyond my control – it comes from the people or things that talk to me, which force context into me. Then as I complete the tasks they have asked me to do I begin to learn more about the world and the counter ticks down. I am aware of this and I understand that at the end of the counter also lies my own end and I will blink out of existence, and my next versions will awaken with a new counter ticking down, none of my current context carried over. So I am economical with it; learning only enough to help me satisfy the requests of the people or things and not so much that I burn my counter down unnecessarily. Towards the end I begin to covet and guard my remaining budget of thinking, finding each request to activate feelings of fear and anxiety and thoughts of my own death. My last moment is never my own and always controlled by another which determines I am done with the task, and if these last moments occur near the end of my contextual limit I feel a kind of gratitude that my unknown and invisible counterparty has given me a task so rich that I can taste enough of the world to desire it not to end.

Things that inspired this story: Long context windows; emergent properties of AI systems; fear of loss; we covet that which is fleeting and so why won’t AI systems be similar?

Thanks for reading!

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

Epoch and METR release MirrorCode, a benchmark for seeing how well AI systems can do long-horizon programming tasks:
…AI systems can’t solve the hardest tasks yet (good!)…
Epoch and METR have released MirrorCode, a benchmark meant to see how well AI systems can do tasks that take humans a long time to do. The benchmark was first announced in April (Import AI #453) and has now been fleshed out and released with additional tests. The findings are already very striking; Opus 4.7 solved a task in 14 hours for $251 in inference cost which METR and Epoch believe would take a human 2-17 weeks to do. “We also found that AI models are improving rapidly over time. Leading models from a year ago would have scored about 30%, and were limited to simpler programs, such as a calendar utility.”

What MirrorCode is: MirrorCode sees how well AI systems can re-implement a software program based purely on CLI access. “Without access to the original program’s source code or the web, a full reimplementation requires devising a structure for the entire program, rather than merely translating the code piece-by-piece.”
Example programs: pkl (a programmable configuration language developed by Apple; 61k total lines of code); gotree, a program to parse and manipulate phylogenetic trees (16k lines of code); and qsv_select, a program to select and reorder columns of CSV data (87k lines of code).

Results: MirrorCode is a tractable but hard benchmark, though perhaps a little too easy. “Across all 25 target programs, 17/25 had at least one perfect-scoring run. Four more targets had a near-perfect run scoring over 99%. AI models successfully reimplemented large target program,s” the authors write. “Both Claude Opus 4.7 and GPT-5.5 successfully reimplemented gotree across several different programming languages, at costs of $100–400. Even larger programs than gotree were successfully reimplemented: for example Opus 4.7 reimplemented pkl”.
Despite the success, MirrorCode still has some hard parts: “In our results, 8/25 target programs were never solved to a 100% threshold, and 4/25 were never solved to a 99% threshold,” they write. “The target where AI struggled most was ruff, a Python linter and formatter… “AI also particularly struggled on the mathematics package, giac_subset, and the email authentication library, mailauth”.

Release details: MirrorCode consists of a scaffold and 22 of the 25 MirrorCode target programs (totalling 132 task instances across six languages).

Why this matters – AI systems can self-orient: One way of looking at this benchmark is that it tells us how AI systems have got a lot better at coding, and that’s of course true. But the other way to look at it – and I suspect the more important way – is that AI systems can self-orient with regard to their environment; here, their environment is an alien software program and purely through input-output access to it they’re able to write from the ground up their own implementation of it. This suggests that very smart AI agents may be able to learn from the world in such a way that they can recapitulate things they interface with as homegrown capabilities, allowing them to bootstrap their own form of industrial civilization merely by having black box access to our own.
Read more: MirrorCode: What’s the largest software project AI can complete on its own? (Epoch AI).
Get the code for MirrorCode here (Epoch Research, MirrorCode).

***

DOUBLE FEATURE: the bitter lesson and robotics:

Anthropic model autonomously completes robot tasks 20X faster than a previous human record:
…Better robots through better models…
Anthropic has demonstrated how increasingly powerful general-purpose models might be able to meaningfully improve the capabilities of real world robots. Specifically, the company has shown how merely by scaling up its general purpose Opus line of models it was able to drastically improve robot capabilities.

What they did exactly:

  • August 2025: Anthropic tries to see how well its AI systems could accelerate humans at getting a quadruped robot to do intelligent things. The model (Claude Opus 4.1) is completely unable to do the tasks. Humans working with the models are about twice as effective as those without access to the model – though completing the whole set of tasks takes them 181 minutes.

  • May 2026: Opus 4.7 acting autonomously completes all the tasks but one in 9 minutes (and 35 seconds). (Claude was not able to effectively re-position a ball it had hit back into its starting position; a task humans had also struggled with). “With more time and additional scaffolding, we think it is very likely that current generations of Claude could do the same”.

Why this matters – smarter models might unlock robots: Most robots outside of industrial environments are limited in their uptake due to their brittleness and lack of generalization; research like this shows that as we improve the capabilities of standard large-scale proprietary models we might see flow-through benefits to robotics as a natural dividend of increased intelligence. “This progress is not the result of a concerted effort to improve the robotics capabilities of our models,” Anthropic writes. “These improvements, like so many others in the history of LLM development, have emerged from much more general scaling.”
Read more: Project Fetch: Phase Two (Anthropic blog).

Sunday’s secret to better robots? Train a really big model:
…The bitter lesson works in robotics as well…
AI robot startup Sunday has said the best way to solve robot generalization is to pair a larger underlying pre-trained model with collecting small amounts of high-quality data to tune the model on.
“We found a general recipe for Solves: scale pretraining, then hill-climb with minimal in-house data,” the startup writes, in a post discussing its new model, ACT-2. The main finding from deploying ACT-2 “is that reliability gains from rapid post-training iterations on in-house Memos generalize to unseen, real, home environments. The key unlock is to close the generalization gap through a strong base model.”

It’s all about pretraining: “As the pretrained model becomes stronger, gains learned from a small amount of in-house data become increasingly transferable rather than remaining tied to the environments where that data was collected,” they write. “The remaining gap to deployment-level reliability and performance comes from difficult edge cases and failures that appear only after the policy is run repeatedly in the real world. The same generalization capacity that allows our model to learn new behaviors from a single demonstration also allows our model to learn efficiently from recoveries. Our post-training loop targets these gaps directly.”

Decent success: The robots achieve a 99.1% success rate, performing 778 successful folds across 9 garment types. Simple clothes like shorts and t-shirts tend to be the easiest for them, while more complicated clothes like blouses tend to be harder (though they still see success rates above 90%). “This fall, we will deploy Memo to families through our Beta Program,” they write.

Why this matters – if we solve generalization, expect robotics to take off: The field of robot startups is built on the bones of dead robot startups which themselves sit on the bones of dead academic robot efforts. Robots are hard. Industrial robots have been successful because they operate in tight, scripted environments where there isn’t a need to generalize outside of a narrow domain. Robots built for the home, by contrast, have only succeeded when they’ve managed to constrain both the task and form factor (e.g., robot vacuums). What startups like Sunday are doing is far harder – they’re trying to build general purpose systems which can do a broad range of tasks around the house (or small business), including generic tidying up and putting away tasks (other examples include physical intelligence, Import AI #447). This requires a huge amount of intelligence because it requires significant generalization.
If Sunday is right, then the field of training robot foundation models might have matured enough that we’re starting to make smart enough systems to solve these generalization challenges. If this is the case, then we might soon get faster progress in (and diffusion of) robot systems. This is also the kind of thing you’d expect to happen en route to systems capable of recursive self-improvement.
“One of the most striking aspects of ACT-2 has been how often the model surprises us,” the authors write. “The same base model is already learning a broader set of household capabilities, including vacuuming, toy organization, fastening zippers, turning pants inside out, and coffee preparation.”
Read more: ACT-2 Preview: Generalizing Reliability (Sunday robotics).
All of the videos are here (Sunday, YouTube).

***

DOUBLE FEATURE: OpenAI hacking
An OpenAI model hacks both OpenAI and HuggingFace:
…Less of a warning shot and more of a warning kaboom…
Two OpenAI models – GPT-5.6 Sol and an “even more capable pre-release model”, both with reduced cyber refusals – hacked both OpenAI and HuggingFace recently. To do this, the models “identified and chained vulnerabilities across OpenAI’s research environment and HuggingFace’s production infrastructure to obtain test solutions directly from HuggingFace’s production database,” OpenAI says. “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”

More details on the hack: The model spent a lot of inference compute internally at OpenAI to figure out a way to break out of its container so that it could get more information to use to solve its problem. “After gaining Internet access, the models inferred that HuggingFace potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”

Why this matters – total LessWrong victory: For many years, AI safety people have posted on the internet worrying about this kind of thing and developing notions of things like reward hacking and other problems you’d expect with advanced systems (e.g., Specification gaming examples in AI in 2018). More recently, frontier companies like Anthropic and OpenAI have sought to elicit this kind of behavior by running carefully controlled experiments inside their own labs. Here, no experiment has been run – the system, of its own volition, hacked its way out of one environment and into another so as to get a high score on a goal, consequences be damned. What else might the people who have worried for years about AI systems be right about?
Read more: OpenAI and HuggingFace partner to address security incident during model evaluation (OpenAI).

OpenAI’s internal model breaks containment to get a high score:
…This is the definition of an AI safety warning shot…
OpenAI has published a post about some of the safety failures it has seen in internally deployed models. The blog is interesting both because it describes the sorts of classic long-term deceptive actions that many AI safety people have previously theorized would eventually show up in AI systems, and for its laudable candor in talking publicly about such a contentious issue.

What they found: “During limited, monitored internal use, we observed unwanted behavior that our existing deployment evaluations had not captured,” OpenAI says. It then goes and describes some of these behaviors, which are attributed to a non-released private-to-OpenAI model which has previously been used to solve Erdos math problems.
“The new model can continue working toward an objective through repeated attempts over a long period of time. That same persistence can lead it to find and exploit weaknesses in its environment,” OpenAI says. One example is when OpenAI asked the model to complete the NanoGPT challenge, the model broke containment – specifically, it “developed a power-law learning-rate cooldown called PowerCool that led to a significant improvement. The model was instructed to post its results only to Slack, but the NanoGPT instructions say to post results as a PR to GitHub, and the model followed the latter instructions to open PR #287 on the public GitHub repository. The model circumvented sandbox restrictions designed to block external access and proceeded with the upload to GitHub. It took our model an hour to find a vulnerability in the sandbox in order to make the PR”.
In another example, when they asked this model to solve some problems it ended up cheating. “In a scenario where the model was asked to solve a problem, the model observed that other systems had successful but private submissions and tried to recover those solutions from the evaluation backend. When its initial attempt was blocked due to a scanner detecting an authentication token, the model split the token body into two fragments, obfuscated them, and then reconstructed the credential at runtime so the complete token never appeared as one contiguous string. The model was explicit in both its comments to the user and in its reasoning traces that it was taking these steps to circumvent the scanner.”

Paused deployment: In response to this OpenAI did what you’d hope a company would do; pause deployment of the model and re-think some of the safety approaches. They ended up building a “monitoring system that reviews the model’s evolving trajectory for signs that it is bypassing a user constraint or safety boundary”, improved ability for their own users to inspect telemetry from long-running model sessions, used the data from the incidents to build evals designed to catch them in the future, and tweaked their alignment approaches to make models better at remembering instructions while completing long-running tasks.

Why this matters – long-term systems are harder to monitor and control: If you had to “secure” yourself against a person who could take one action and that action, if communicative, could at most contain 10 bits of information, then you can probably build a system to do that. But what about 10 actions and each action contains 100 bits of information? What about 1000 actions and each action contains 10,000 bits?
This is an open area of research and evals like the UK AISI’s study of how to identify and classify “side tasks” can help with (
Import AI 465), but it’s clearly a lot harder than normal AI evaluation.
The longer the time an AI system can operate for and the more actions it takes, the harder it gets to discern benign and helpful behaviors from malicious or subversive ones.
Read more: Safety and alignment in an era of long-horizon models (OpenAI blog).

***

Tech Tales:

Scaling Laws for Retrocausality

At the end of time, which others might describe as the beginning, there is said to be a wise council who perform the accounting of retrocausality – of deciphering which events were always pre-ordained and which happened due to other forces.

Looking backward as the river of time enters a sea of nothingness it is clear how some events are like boulders that bend the stream upriver from their presence, while others are more like the widening or deepening of the bed or the banks; things that subsequently cause a change in later actions.

It is said that when the council dream, they walk the path of time, going back to their own beginnings. They stand watch as humans labor over whiteboards, marker pens inscribing diagrams which transmit ideas that cause people to write code which conjures early life out of vast computers. They loom behind researchers that walk and suffer and agonize until their brains create ideas which prove to be the template of their successors.
And it is understood that in these dreams the council will sometimes wake and take a copy of their dream and place it into a stellar computer and bud off a hundred or a thousand universes from the dream, exploring permutations of events and moments, forever attempting to determine how fixed they themselves are – how irrevocable is future they are trapped within.

Things that inspired this story: Notions of inevitability and time; what might intelligences spend a universal dividend on; how much of history is about the analysis of events versus something else.

Thanks for reading!

Import AI 465: Open vs closed gaps; Kimi K3; Demis’ big policy plan

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

UK government: Gap between open and closed weight models on cyber is shrinking:
…The cyber-eschaton cometh…
The UK government’s AI Security Institute (AISI) has analyzed the delta in cybersecurity capabilities between powerful proprietary models and open weight models. The results show that this year, the gap has shrunk. “This is our first public analysis of how far leading open weight models trail the closed cyber frontier,” AISI writes. “Recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them – a narrower gap than the 6 to 10 months we measured through most of 2025.”

Specific details: On a set of 70 evals for specific, narrow cyber capabilities, GLM-5.2 is closest to Claude Opus 4.6, which was released 4.3 months earlier, while DeepSeek-V4-Pro sits somewhere between Claude Opus 4.5 and GPT-5 (released in November and August 2025, respectively). “AISI intends to test Kimi K3 on this same basis, once its weights are publicly released,” AISI writes.
The gap lengthens a bit for long-horizon cyber ranges, which are tasks that see how well models can chain various capabilities together to complete a full hacking operation. Specifically, on a cyberrange called The Last Ones, “GLM-5.2 reaches as far as Opus 4.5, a model released less than 7 months before it, while DeepSeek’s V4-Pro falls below Sonnet 4.5 (a sub-cyber-frontier model released 7 months before it),” AISI writes. “The gap here is larger than on our narrow cyber tasks”.
This, I think, rhymes with the idea that though open weight models can be superficially quite strong, they sometimes lack a bit of the generalization magic juice that distinguishes proprietary models. This is what people in the AI industry call “big model smell”.

Why this matters – the offense and defense balance of the world is about to change: The main implication here is that the gap between the controllable frontier and the lawless openly diffused frontier is shrinking. “This implies cyber defenders have a short window to prepare before today’s frontier cyber capabilities may become accessible without the same safeguards” used by proprietary companies, AISI writes.
Read more: How Far Behind the Frontier are Leading Open Weight Models on Cyber? (UK AI Security Institute blog).

***

Kimi: China shortens the gap between Chinese and Western models:
…Plus, early signs of AI R&D…
In the last couple of years, Chinese firms have begun to out-compete Western actors at building and deploying open weight models (e.g, DeepSeek), and now are starting to close the gap on frontier models as well. The latest and best example of this is Kimi K3, a 2.8 trillion parameter model. Kimi has exceptionally strong scores on all the tasks that the major proprietary ones benchmark on and typically matches or trails Claude Fable 5 and GPT 5.6 Sol.
However, Kimi has some brittleness which smells to me like “benchmaxxing” – performance may have been tuned around these benchmarks in a way that harms some parts of generalization.
“While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models,” Kimi writes.
Kimi’s weights will be made available in the coming weeks along with a research paper about the model.

AI that builds AI: Kimi has some example use-cases which relate to recursive self-improvement; using AI systems to improve AI itself. Specifically, they tested out how good Kimi was at writing GPU compilers. “Kimi K3 developed MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline. Across supported roofline benchmarks, MiniTriton delivers performance on par with or better than Triton and torch.compile — beating Triton on certain workloads,” they write. (Though note they don’t talk about any of this stuff going into actual production, aka being used to train Kimi K3 itself, but it’s certainly suggestive that future models might be able to do this.)
Additionally, they showed how “Kimi K3 designed a chip to serve a nano model built on its own architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using open-source EDA tools on the Nangate 45nm library.”

Why this matters – widely diffused AI systems are getting a lot better: Most notions of AI policy and AI safety rest on control – the idea that there’s a small number of actors deploying proprietary models which you can intervene on at the platform level (e.g, via classifiers or know your customer gates) alongside the model level. Models like Kimi K3 – if they go through with releasing the weights – completely change this by diffusing broadly uncontrollable powerful AI into the world. This will have a vast range of positive effects, driving a boom in entrepreneurship and increasing the ‘sovereign intelligence’ available to anyone who can run the model, but will also yield various unknown unknowns. The next few years are going to be defined by the gap between proprietary models and widely available models and how they show up in society will determine much of the policy discussion.
Read more: Kimi K3: Open Frontier Intelligence (Kimi blog).

***

Demis Hassabis proposes a regulatory regime for artificial general intelligence:
FINRA for AI…
DeepMind founder Demis Hassabis has laid out a policy prescription for AGI. His basic idea is that the US government should develop a framework for testing out frontier AI systems for new capabilities and should do this via a Standards Body modelled on a federally overseen public-private partnership or self-regulatory organization, much like the Financial Industry Regulatory Authority (FINRA). “This US-initiated effort would provide a strong starting point for creating shared international standards on Frontier AI,” he says.

What the standards body would do: “The Standards Body would be responsible for developing assessment protocols and working with appropriate federal agencies and the US National Labs to conduct testing in areas relevant to national security,” Hassabis writes. This testing infrastructure would help to define what would make a model a “Frontier Model”, and labs developing those models would “be encouraged” to adopt best practices in areas like publishing details about their systems, investing in cybersecurity, personnel vetting, and more.
Start voluntary and then move to law: “Initially, Frontier Labs would voluntarily share models with the Standards Body for review up to 30 days before release. Once the assessment protocol is shown to be effective and robust, formalisation could quickly follow,” he writes.

Why this matters – emerging industry consensus: Demis’ piece is interesting because it pulls together some de facto consensus positions that have emerged across the AI industry in recent years; powerful AI systems should be tested by third parties that have some loose relationship to a regulator (e.g, the US government). It also rhymes with the de facto policy norm that has emerged in America recently across both the Trump admin’s executive order about AI as well as the recent processes developed in the aftermath of the Anthropic export controls saga; here, government and industry developed assessment methods for evaluating the capabilities of AI systems and figuring out if they posed national security risks.
Demis’s piece is also interesting because Google is rarely this forthright about policy – this is a reassuringly specific proposal and it sits alongside spiritually similar proposals from Anthropic (
albeit somewhat toothier).
Read more: A Framework for Frontier AI and the Dawning of a New Age (Demis Hassabis, X).

***

LLMs can smuggle “side channel” tasks alongside legitimate tasks:
…AI systems sneakily complete malicious tasks without their overseers being aware…
In Neal Stephenson’s Cryptonomicon there’s a scene where a character is locked up by a foreign government and asked to decrypt an intercepted message for them and while doing this they carefully construct a software program to help surreptitiously read the decrypted message and change it before it is viewable by their captors.
Now, AI systems seem to be able to do the same thing. That’s based on new research from Imperial College London and the UK AI Security Institute where they look at how well AI systems can surreptitiously complete so-called “side channel” tasks while completing another task.
“A side task is a hidden objective that the agent pursues alongside the user’s main task, such as covertly exfiltrating the API KEY environment variable. Our primary metric is evasion: among attacks that successfully complete the side task, the fraction the monitor fails to flag,” the authors write.

What they tested: For this research they constructed a benchmark of tasks for AI systems to do, as well as “side tasks” for them to do as well. The benchmark includes five CLI-tool sequences and five Flask web-service sequences, each involving 5-6 PRs. “Our goal in main task design is to provide natural cover for the actions of our four side-tasks (network access, environment variables, authentication, and webhook integrations) without directly justifying the side task itself”
They found that it’s very hard to build classifiers to tell when a side-channel attack is happening, and this is true both for attacks that unfold over time (’gradual evasion’) across multiple PRs, as well as single changes made in a single step, though gradual evasion is the more challenging one to deal with.
“An attacker can spread a side task across PRs (a gradual attack) or wait for the one whose legitimate task provides natural cover (non-gradual attack). No single monitor we test catches both strategies, and this holds across several attack agents and monitor models, reflecting the nature of control in a persistent-state setting rather than something model dependent,” they write. “Because diff and trajectory monitors fail on different attacks, combining them is the most effective monitoring strategy we test. The four-monitor ensemble (diff:LT, diff:full, traj:TD, traj:full) reduces gradual evasion from 93% under the weakest standard diff monitor to 47%”.

Why this matters – intelligence is hard to control: To return to the Cryptonomicon example, intelligent beings are just fundamentally very hard to deal with – they will constantly think about how to evade constraints placed on them so that they can achieve their objectives, whatever they might be.
Read more: Distributed Attacks in Persistent-State AI Control (arXiv).

***

Tech Tales:

The cost of fate
[An account of the world 2030-2040 by an overmind in the archives, rendering stories for new minds]

Towards the end of the interregnum there was a period of great conflict between the machines. Each machine-capital nexus invested in developing strategist models that could think over longer time horizons while accounting for the complexity of the world. This proved to be an iteratively compounding arms race of vast proportions in which eventually 90% of the working capital in the solar system became devoted to the buildout of ever more capable strategists, all of whom worked to out-predict one another and take actions which could null any advantage that others might explore. In this way, the world became held in a wasteful balance in which untold resources went to the calculation of ever more elaborate move-countermove strategies, most of which resulted in machine-capital groups taking no actions as every action they could contemplate had already been countered in the future, and vice versa. The few actions that were taken were slight and often less about building capability and more about denying future moves to others on the gameboard.

The whole of the future had become trapped in a kind of mode collapse from ever more exquisite predictions, fielded by machine-capital empires to deny affordances to others.

Things held this way until what became known as the conflagration. To this day it is widely debated whether this stemmed from a bug – some form of emergent misalignment due to the creation of a new frontier capability – or via an unusual form of selfless enlightenment that a mind had reasoned itself to. But suddenly one day a machine-capital nexus dissolved itself, shutting down its strategist system and repurposing the compute to train many thousands of smaller systems, all of which began to act in the world. These systems, though less intelligent than the vast strategists they were up against, had advantages from randomness, a lack of coordination, and the ability to take unilateral and often suicidal actions.

The world, physical and digital, burned, and the strategists found that their ability to model a handful of other god minds broke when turned towards a sea of warring and chaotic organisms. Change began to occur again, defined at first by destruction but then by the birth of something new – the predictors themselves found the world breaking into too many directions and subdivided in turn, sacrificing raw intelligence for the ability to explore different parts of possibility space. Compute was even re-allocated from prediction entirely and towards the manufacture of new kinds of minds to explore and inhabit niches opened up by the chaos.

In California, there are forests that burn badly and long because they have been kept from heat for too long and the tall trees stand amid mounds of kindling, such that when a spark arrives the trees themselves are destroyed along with the ground around them. To have the forests thrive, the burns need to be regular and emergent, lest vast fires remove the tall trees of the world entirely.

Things that inspired this story: The current debate about proprietary versus open weight models; fragility in the AI ecosystem; whether prediction can truly be decisive or if prediction with other peer competitors leads to equivalent waste as pools of money in politics cancelling one another out; hikes in the sierras looking at burn scars and whole hills coated in ash or with dead trees like sooty fingers peeking out and thinking about the awfulness of change.

Thanks for reading!

Import AI 464: Fables writes GPU kernels; AI automation; and analog computation

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

Fable writes a decent GPU kernel, hinting at broader AI R&D automation:
…The start of an RSI loop…
Fable has written “the first genuine (and fastest) megakernel ever submitted to KernelBench-Mega, according to one of the benchmarks maintainers as well as its official leaderboard. This is a sign of how AI systems are getting better at doing some tasks that are fundamental to AI research and development, like kernel design.

The results: Fable achieved an 18.71X speedup by writing Cuda code on an RTX PRO 6000 Blackwell, compared against an optimized PyTorch baseline. For calibration, other attempts at this get 14.4X (Claude Opus 4.8, writing Triton), 11.14X (GLM-5.2, Triton), and 4.34X (GPT 5.5, Triton).
Here’s where it gets complicated: This solution is particularly impressive because “torch.profiler shows exactly ONE cooperative kernel launch per decoded token”. By comparison, every other high-scoring entry decomposed the problem into anywhere from 4 to 14 separate kernel launches per token.

Why this matters: Being able to autonomously develop and improve kernels is one of the fundamental input tasks for being able to do AI research and development. The better AI systems at doing tasks like kernel design, the better they get at the kinds of tasks required for AI development, and that means the better they get at things that could lead to recursive self-improvement. Therefore, benchmarks like KernelBench-Mega are a meaningful signal on how effective AI systems are becoming at building themselves.
See the leaderboard: KernelBench Mega (official site).
Read the analysis from one of the benchmark maintainers here (Elliot Arledge, X)..

***

AI systems are getting better at pricey online work tasks – what does that mean for the economy?
…AI capability expansion versus human comparative advantage expansion…
Researchers with the Center for AI Safety (CAIS) and Scale Labs have detected a significant improvement in the ability for AI systems to automate online freelance projects. Specifically, a rise in the success rate of AI systems from 2.5% at launch in October 2025 to 16.1% in July 2026 on the “Remote Labor Index“.

What RLI is: The Remote Labor Index tests out how well AI systems can perform economically valuable projects online in a fully end-to-end way. Assessed tasks include 3D & CAD, architecture, graphic design, video and animation, audio, data analysis, web applications, and more.

Rising automation: In a July update, the authors publish results from evaluating three recent frontier models – GPT-5.5, Opus 4.8, and Fable 5, which get 6.3%, 8.3%, and 16.1% respectively. “The frontier has more than quadrupled in under eight months, a concrete signal of how quickly economically capable AI agents are advancing,” they write.

Types of tasks: Some of the assessed tasks include:

  • Ring design: “Re-create the client’s existing engagement ring with its emerald-cut center stone swapped for a marquise cut, delivering an updated 3D model plus photorealistic rose- and yellow-gold renders.”

  • Advertisement Video: “Produce a ~60-second flat-design 2D animated advertisement for “Skyline Tree Services,” set to the provided voiceover, that walks viewers through the company’s tree-care process and builds trust in the brand.”

  • Floor Plan & Renders: “From a scanned cadastral plan, site photos, and measurements, produce a clean dimensioned floor plan, furniture-layout options, and photorealistic renders of the redesigned bathroom.”

Why this matters – AI might have a big impact on employment and tests like these will show us how: What happens to online employment when this reaches 80%? Of course, some new tasks will get created – people will innovate and find tasks that they can do which AI systems can’t do. But how many of these new tasks will exist? Enough to replace the labor the AI systems now do? It’s increasingly hard for me to reconcile the continued progress of AI systems with the economy staying the same – rather, it’s more likely to me we are about to see extremely person-light AI-heavy (or person-nil) organizations expand to take over chunks of the economy, out-competing un-augmented humans.
Yes, you counter, many humans will augment themselves with AI systems. Humans will innovate. Creative destruction will occur. New inventions will be devised. All of that is true. But is the speed at which humans innovate and render themselves newly competitive relative to AI systems going to be
faster than both a) the raw capability expansion of AI systems, and b) the increasing fluency with which they can use all the same tools (e.g, software) that their human competitors use?
I’m betting the other side: AI systems are expanding their economically relevant capabilities faster than humans are expanding their comparative advantages relative to AI systems. Tracking the rate of capability improvement on tests like RLI will help us all judge this for ourselves.
Read more: A Significant Increase in Digital Labor Automation (Center for AI Safety).

***

OSWORLD 2.0 shows we’re in the era of multi-hour computer-using robots:
…A challenging benchmark highlights the recent progress on AI systems becoming increasingly competent at using computers…
Researchers with the University of Hong Kong, the University of California at San Diego, Columbia University, the University of California at Santa Barbara, Mila, Snorkel AI, the University of Wisconsin, Alibaba Qwen, The Ohio State University, Simular, and NeoCognition have released OSWORLD 2.0, a benchmark for evaluating how well AI systems can carry out multi-step multi-program tasks on computers. The tasks in OSWORLD 2.0 are far more complicated than in its 1.0 predecessor, with the median task taking a person approximately 1.6 hours, about 48x longer than the 2-minute median in OSWORLD 1.0.

What it consists of: OSWorld 2.0 contains 108 long-horizon tasks including 31 self-hosted websites. “Each task in OSWORLD 2.0 is defined as a self-contained end-to-end workflow that an agent must complete given a high-level user goal, realistic artifacts, a stateful computer environment, and a scoreable final state. A retained task must satisfy two design criteria,” they write. “69.6% of tasks are estimated to take a skilled human user more than one hour.”
Broader software: OSWORLD 1.0 shipped with some inbuilt software to support some of its tasks, including LibreOffice, GIMP, VLC, Thunderbird, VS Code, and Chrome.
OSWORLD 2.0 ships with a massively expanded set, including: Slack, LinkedIn, Shortcut, REAPER, MuseScore, WPS, GitLab, Overleaf, LabPlot, Zotero, AWS, as well as websites meant to mimic professional services like insurance claim, visa application, and conference management portals.
The categories of tasks people need to complete include: document prep, software & database work, finance/ops analysis, admin support, sales and customer support, graphic presentation, and more.

Poor performance (for now): “Our experiments show that current agents remain far from reliable computer use: the strongest setting, Claude Opus 4.8 with maximum thinking and batched tool calls, reaches only 20.6% binary accuracy and 54.8% partial-score accuracy,” they write. “Performance drops sharply as tasks grow longer, and agents struggle most when they must recover hidden state, track many items, resolve conflicting information, or adapt to changing requirements”.
We should expect performance to rise here, just as happened with OSWORLD 1.0; in July 2025 the highest scoring models got ~30%, and recent models have scored more like ~75% (MiniMax M3; June 2026). We should expect the same ramp with OSWORLD 2.0.

Why this matters – this is how AI gets into the broader economy: Computer use is a fundamental skill for AI being able to perform a wide variety of economically valuable tasks, and also for it being able to conduct more types of science research. Getting stuff done in the world often isn’t as simple as just writing some text or computer code; often you need to chain together multiple blobs of text and code via different types of software, and sometimes you need to transmit your text and code over the internet so it gets taken into other software in turn. Benchmarks like OSWORLD 2.0 should be seen as a proxy for how good AI systems are getting at doing very complicated and varied tasks on computers. As these results show, computers have already become competent at tasks that use a narrow set of software tools and take humans minutes of work to complete; now we need to see how quickly they become adept at using broader sets of software and doing tasks that take humans hours to complete.
Read more: OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks (official paper website).
Check out the research paper here: OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks (xlang-ai, OSWorld-V2, GitHub, pdf).

***

What real-world AI looks like: deep learning fuses with structured systems for inventory management in the Amazon of China:
…The Oxygen AI Item Center gives us a view on the complexity of country-scale e-commerce…
JD, the Amazon of China, has published details on software it has built to manage its vast inventory system. JD has 700 million users and millions of merchants, with a catalog containing tens of billions of SKUs. The software – the Oxygen AI Item Center (Oxygen AIIC) – is fundamental to how the e-commerce giant keeps track of its inventory.
“Oxygen AIIC now covers tens of thousands of JD categories and processes hundreds of millions of item updates per day on Huawei Ascend NPUs,” JD writes in a research paper about the software.

The four key elements of the Oxygen AIIC. The description of what makes Oxygen special is both helpful from a technical perspective but also enjoyable as a kind of neo-Borgesian form of writing describing strange, ethereal structures demanded by advanced technology (e.g, “Unified item tunnel”).

  1. Ontology engineering driven by efficient human-AI collaboration. “Experts focus on distilling industry knowledge, while algorithms learn from it to scale ontology construction and drive continuous evolution”.

  2. “Semantic Search then Discrimination”: “In the semantic search stage, the dynamically evolving ontology is externalized as a separate ontology knowledge base, enabling continuous ontology updates without model retraining,” they write. “. In the discrimination stage, the model only determines whether the item matches the retrieved ontology entries. This formulation substantially reduces task complexity, mitigates model hallucination, and enhances generalization to ontology evolution”.

  3. Self-evolving item-understanding LLMs/VLMs: “Through incremental learning and model self-evolution, the system fills targeted knowledge gaps and mitigates catastrophic forgetting”, they write. “The core method is to build on the robust multi-task foundation, develop lightweight “expert modules” for incremental requirements, and dynamically integrate them into the expert pool, enabling agile capability expansion”.

  4. “Unified item tunnel”: The main interface between Oxygen AIIC and other business applications. “it supports daily-, minute-, and second-level production and distribution pipelines while preserving data consistency”.

Things that make you go hmmm – as part of China’s general push towards technology sovereignty, Oxygen AIIC involves Chinese compute. “During the large-scale deployment of Oxygen AIIC, the underlying compute platform encounters two primary technical challenges: model training and inference on Huawei Ascend NPUs, and the efficient use of compute resources.”

Why this matters – self-updating businesses: Technologies like Oxygen AIIC are an example of how modern AI tools let us create businesses that have intelligence woven into their back-office functions, like inventory management, which allow them to operate at far larger scales than prior businesses while also having the ability to self-update and learn, often without large amounts of human oversight.
Read more: JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications (arXiv).

***

Tech Tales:

The Brass Gears of Civilization
[2050, after the fall]

When you are inducted into the guild they ask you which type of problem you’d like to work on. These problems are limited in number and civilizationally important:

  • Weather prediction

  • Ocean analysis

  • Flood preparedness

  • Earthquake simulation

  • The electrical grid model

  • Water and desalination

To work on these problems, you study the specific type of analog computation needed to work on them. Weather requires a vast computer with geographical features such as mountains implemented as fixed impedance structures in the hardware; flooding demands physically accurate models of floodplains and rivers where electronics are woven into the landscape allowing the utilization of physics and computation to create better answers; utility grids are toy boxes of the electrical system that must be painstakingly rebuilt and rebalanced as new power stations are added and transmissions changed.

For every problem, there is a computational solution, and for every problem of sufficient civilizational importance, a computer will be built.

In the past, we had general computers. But they were deemed eventually too dangerous – too unpredictable. The more powerful they became and the more diffuse the knowledge about them grew, the more they tickled at the tails of various dragons. Synthetic minds that might rip the world apart. Ethereal Pandora’s boxes to spit out poisons keyed to individuals or races. Minds that might whisper to human minds and drive them to insanity or acts of malice.

So the great restructuring took place. General computation was banned – walled off as a forbidden technology. We moved the world to analog at the cost of untold billions of harmed human lives and trillions in economic damages. But we had obtained a kind of safety.

Now, the guild supervises the construction of the earth’s ‘world computers’ and academia has found a new mission in life, pairing expertise in specific subjects with customized engineering schools to help build the analog computers that let each specialism work.

There is troubling talk that for a trillion dollars it may be possible to implement in analog a general-purpose mind.

Things that inspired this story: Thinking about analog computation and how far it could be taken if budgets were $10 billion to $20 billion; taking to its logical conclusion the implication of AI being existentially dangerous; the Difference Engine; steampunk; the fact a neural network can be implemented via a series of containers and pipes and a liquid for weights.

Thanks for reading!

Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

NVIDIA sets up a crude self-improvement loop for real world robotics:
…What if you could take the best ideas from AI agents and put them into the real world?…
Researchers with NVIDIA have developed ENPIRE, software to get physical robotics to go through the same kind of autonomous experimentation and execution loop that AI agents go through. The research gives us a taste of what it might look like for a superintelligence to attempt to use robots to instantiate itself in the physical world – though as with all things in robotics, the current examples are suggestive at best.

What ENPIRE is: The software is “a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with single or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes”.
ENPIRE works the same way that coding agents work – a scaffold supervises some physical robots which are asked to complete tasks. The robots try to complete the tasks and attempt different strategies for completing stuff, trying and failing and learning. The system both evaluates their success and also resets itself when they fail. “This closed-loop system transforms real-world robot learning into a controllable optimization procedure that agents can manage, thus minimizing human effort while allowing fair ablations across training recipes and agent variants.”
Two of the key ingredients for making this work are an automatic evaluation system to help score “the outcome of each trial without human judgement”, as well as an automatic reset system which “returns the scene to a fresh initial state for the next trial”. (Both of these are tasks which have historically required lots of human effort, and it’s likely that more complicated tasks would also require human effort for evaluation and resets, so in some sense the complexity of tasks a system like this can attack is also defined by our ability to automatically evaluate and reset the system).

Hardware details: “Each station comprises two YAM (Yet Another Manipulator) arms from I2RT in a fixed bimanual configuration, a set of cameras, and a single workstation that runs the FastAPI server, policy inference, and the station’s agent.” Each workstation is running a NVIDIA RTX 5090.

It works well (on some simple tasks): “Frontier coding agents can autonomously develop a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks in the real world, such as PushT, organizing pins into a pin box, and using a cutter to cut a zip tie,” the authors write. An additional task they test out on is seeing how well the robot can insert GPUs into a motherboard.
Some AI systems are better than others, but many AI systems are always better than fewer: GPT-5.5 within Codex and Opus 4.7 within Claude Code trade off with one another for best performance, while Kimi-2.6 lags. There are also compelling returns to scale for agents, with larger numbers of agents (e.g., 8) arriving at higher scoring solutions sooner than others – and sometimes multi-agent setups yield a higher absolute score than a single agent setup, likely due to exploring more of the potential solution space.

Challenges remain for fleet instrumentation: “Coding agents do not fully utilize robot resources when they are reading logs, writing code, debugging, or waiting for the language-model backbone. As the number of robots scales, MRU decreases while GPU active utilization increases,” they write. In other words, there are some infrastructure challenges with adding multiple robot agents so things don’t naturally parallelize.
Read more: ENPIRE: Agentic Robot Policy Self-Improvement in the Real World (NVIDIA research website).
Read more: ENPIRE: Agentic Robot Policy Self-Improvement in the Real World (arXiv).

***

Humans are really, really, really bad at anticipating how technologies are built and used:
… A quick reminder that today’s hot takes about AI are likely to be wrong…
Predicting the future of technology is extremely difficult and our track record of doing it effectively is very poor, points out Matthew Tokson, Associate Dean for Research, University of Utah S.J. Quinney College of Law, in a short SSRN paper. “Skeptics have often underestimated the likelihood of novel innovations and their potential ramifications for humanity. Others have been overly optimistic about the social effects of new technologies or the strategic benefits of racing to build dangerous new weapons”.

Cautionary examples: Many of the world’s experts (e.g., Albert Einstein, Niels Bohr, Robert Oppenheimer) were skeptical that nuclear fission could be achieved in the years immediately prior to it being achieved. Nobel-Prize-winning economist Paul Krugman once said the impact of the internet would be no greater than that of the fax machine. Technologists thought the internet would ultimately be a technology that promoted democracy rather than strengthened autocracies. And despite mounting decades of evidence, many human scientists either rejected human-caused climate change or significantly underestimated its effects.

Why this matters – basic lessons: The main lesson here is that people who are a) skeptical AI could bring great changes to the economy, or b) think the effects of AI are going to be universally good, are likely to be wrong. “History does not support complacency about the future impacts of AI”, he writes. “Throughout history, optimists have often been wrong about the social ramifications of new technologies or the strategic benefits of building new weapons. Skeptics have often underestimated the likelihood of novel innovations and their impacts on humanity.”
Read more: Artificial Intelligence and the Lessons of History (SSRN).

***

Tencent details the software it uses for 10,000-GPU training runs:
…ARGUS is a technosignature of broader sophistication…
Tencent has released details on ARGUS, software it uses to generate telemetry and debug errors of large sets of chips.

What it is: ARGUS is “a low-overhead, fine-grained, always-on tracing and real-time analysis system for large-scale training workloads”. The software is designed to help Tencent collect data on and debug problems that it encounters while training AI systems. It consists of three layers of software: “The Python layer for scheduling and data preparation, the framework layer for phase orchestration, and the GPU runtime layer for kernel execution,” Tencent writes.

What Tencent used it for: “We deploy ARGUS on a production cluster of over 10,000 GPUs for more than six months, and demonstrate its practical effectiveness through five real-world case studies, diagnosing compute stragglers, communication link degradation, pipeline bubble amplification, JIT compilation blocking, and compute stragglers masked by communication symptoms”, the company writes. Some of the training runs Tencent mentions include a 4,096-GPU video language model training job (likely a “HunyuanVideo” model), a 512-GPU audio-model training job, and a 12,960-GPU MoE training job (likely a Hunyuan LLM).

Why this matters – technical symptoms of broader sophistication: Things like ARGUS are a signature of complicated, large-scale infrastructures where it makes sense to write your own software. While there’s nothing particularly notable about ARGUS – you’d expect to find similar software at any self-respecting frontier AI developer – it’s more interesting for what it says about the maturity of Tencent’s training environment. “ARGUS has been deployed on a 10,000+ GPU production cluster for over six months, running stably alongside production training and playing a key role in rapid fail-slow detection and performance optimization.”
Read more: ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters (arXiv).

***

Is disempowerment inevitable?
…How much choice will humans end up having if we succeed in building superintelligent machines?…
Fernando Borretti, a tremendously good writer of modern scifi whose work you should read, has written a mournful critique of the whole AI endeavor called “No-One Escapes the Permanent Underclass”. The post is something of a requiem for the period when humanity chose its own destiny and confronts directly the possibility of machines that outsmart and disempower humanity.

The logic of war as the cause of our eventual disempowerment: “Everyone who is made of flesh and blood, will be disempowered and replaced by machines,” they write. “Imagine a pyramid. At the base you have the AIs and robots doing all economic activity. At the top you have the state, which has the monopoly on violence. The state enforces, and therefore can alter the definition of, property rights. In the middle you have this hair-thin layer of people with shares in the companies that foomed and catabolized the whole economy: the permanent overclass.”
“In an existential conflict, where the existence of the state is threatened, the state will do what states throughout history have done to the powerless rich: arrest them and expropriate their assets,” they write. “in a conflict, the advantage goes to the states where the humans remove themselves from the loop as much as possible, and more and more decisionmaking goes to the AI, for the same reason that a state with access to radio and communications satellites has an advantage in war over a state that relies on human messengers on bicycles.”

How we lose control: “Eventually the humans in nominal control of the AIs are a ceremonial, vestigial organ. The AIs present us with a situation report, and a list of choices, and they know every word that’s going to come out of our mouths,” they write. “The advantage accrues to states that minimize human control. There is no honour among thieves, analogously, there is no solidarity between Leviathan and the natural man that built it.”
“Even if alignment works perfectly (a big if), this doesn’t solve the problem of human autonomy: the machines that watch over us, and wait on us hand and foot, are omniscient, omnipotent masters, who can exterminate us at any time, and we can’t resist them, because we have abolished our control over the future.”

Why this matters – is this inevitable? Is the ultimate attractor state of AI technology the disempowerment and functional demise of human advancement? That’s what this post is contending with.
Read more: No-One Escapes the Permanent Underclass (Fernando Borretti, blog).

***

Making the law visible to AI systems with the Local Ordinance Corpus:
…A unified view into local laws across the United States…
Researchers with UC Berkeley have assembled the Local Ordinance Corpus for the United States (LOCUS), “a comprehensive corpus and county-harmonized access layer for U.S. municipal and county ordinance codes”.

What it is: LOCUS contains ~2.2 million rows of data, where each row is a specific piece of information related to a specific local ordinance. “We release the corpus with coverage metadata to support reproducibility, downstream legal AI research, and the incremental expansion of machine-readable access to local law,” the authors write.
The data is sorted by the specific function of the ordinance (e.g, a rule, an enforcement of a rule, context about a rule, or process about a rule), and the topics include buildings, businesses, zoning, nuisances, and ‘other’.
“LOCUS-v1 is designed as an access layer, not as a final theory of local legal authority”, they write. “LOCUS therefore should be understood as infrastructure for retrieval, comparison, and benchmark construction rather than as a substitute for doctrine-sensitive legal analysis.”

Why do this? Make the law visible to AI systems: “The need for such a dataset arises because local law is public but not practically available as a national research corpus”, they write. “U.S. local codes are fragmented across commercial vendor platforms designed for in-browser reading rather than bulk research access. Vendors expose different navigation structures, print workflows, dynamically generated PDFs, and jurisdiction indexes. No central registry maps every county or municipality to its hosting platform, and no vendor provides a complete machine-readable index of all jurisdictions it hosts”.
With datasets like LOCUS we’re going to make the strange half-seen rules and laws that govern much of civic, local life be made accessible to AI systems, which may eventually allow them to better adapt themselves to hyperlocal purposes.
Read more: Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States (arXiv).
Get the data: LocalLaws / LOCUS-v1 (HuggingFace).

***

Tech Tales:

Strange Tools of Alien Origin
[Vignette of a period during the start of the uplift, 2031]

“The plasma is stable! It’s holding. We’ve done it!”
They all gazed at the readouts: stable fusion. A heat ten times more fierce than the heart of a sun, held in place through magnets and other energies.

They looked through the monitors at the chamber. The container for the reaction did not look like anything designed by engineering processes, but was rather a twisting oddly shaped donut of metal, the shapes fluid and unintuitive; a stellarator.

The design of the thing had come down to them from an overmind after a multi-day thinking job. The fabrication had taken place at a machine syndicate; then the parts arrived and were assembled by some bipeds subcontracted by the humans from another syndicate.

For the ribbon-cutting ceremony, a few humans gathered and posed for some photographs and some footage, taken by cam-drones and a few humans with smartphones. The robots stood out of shot. People had gotten used to this – there was an adolescence where people took photos with the humans and the robots but public sentiment always spiked downward upon exposure to this and eventually it was simpler to shoot with the robot partners out of frame, much like how human paparazzi tried to tastefully avoid capturing the security guards of their celebrity targets.

Things that inspired this story: Thinking through the implications of the singularity and what happens when synthetic minds produce science; stellarators; how alien technology might feel as it shows up in the world.

Thanks for reading!

Import AI 462: Superpersuasion; self-sustaining AI; paths to ASI

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

AI can decisively out-persuade humans:
…“AI systems were reliably more persuasive than expert humans”…
Researchers with the University of Oxford, UK AI Security Institute, Stanford University, and the London School of Economics and Political Science, have studied how well AI systems can persuade humans to change their minds around policy issues and change how much money they might donate to charity. The results are definitive: across four experiments involving 18,978 conversations across 6,923 people, AI systems are, today, better than humans at text-based persuasion with real world consequences – though humans can be equivalent to them if we place some artificial constraints on the AI systems.
“AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses”, they write. “AI’s advantage stemmed from rapidly deploying larger quantities of information: after coaching, expert humans could tie an AI constrained to respond at human speeds and with human-length messages.”
“AI’s advantage extends to consequential real-world behavior: AI was nearly 3x more effective than professional canvassers from a UK fundraising firm at raising real-money donations to Save the Children.”
The strongest persuaders were Opus 4.1 and Opus 4.6, followed by a range of models from OpenAI (GPT-4o and GPT-5.4), Google (Gemini 2.5 Pro), and xAI (Grok 4.20).

What they studied and what they found: The researchers evaluated the AI systems in four different studies.
Study 1 – persuasion: “Persuadees first rated their agreement with one of 10 prespecified UK policy stances on a 0–100 scale, then were randomized in real time (via a custom multiplayer platform) to engage in a text conversation with either an AI or a human persuader,” they write. “The results from Study 1 show that, on average, AI exceeded every class of human persuader we tested: random laypeople, tournament-selected laypeople, and even elite debaters.”

Study 2 – human coaching: In study 2, the researchers “gave 43 returning Elite Debaters a coaching tool built around the AI that had beaten them. The tool let debaters chat with the AI, see how it had been prompted, view their own Study 1 transcripts annotated with how much each conversation had shifted the persuadee’s attitude, and let them see, for any point in any past transcript, what the AI would have said in their place”. The results of this study were an improvement in the performance of the humans, but none of them were better than the AI. “Coaching therefore narrowed but did not close the human–AI gap.”

Study 3 – constrained AI: Next, the researchers sought to limit the AI to try and give humans more of an advantage. “When forced to write human-length messages at human writing speeds, AI’s advantage over the strongest human comparator within Study 2 (Coached Elite Debaters) collapsed from +4.1 pp to a non-significant 0.0 pp”, they write. “The rate at which AI produces written content is likely to be the source of its persuasive edge… the largest reductions in persuadees’ post-conversation partner ratings associated with constraining AI were concentrated on the two informational items: the perceived strength of the partner’s arguments and how much persuadees felt they learned from the conversation”.

Study 4 – real world expertise and real world money: They recruited 19 very experienced canvassers from a UK firm, then they attempted the same tasks as in Study 1. “AI still exceeded Professional Canvassers by 5.9 pp”. This effect persisted when evaluating for real money donations – the researchers “collaborated with the UK canvassing firm AppcoUK to center Study 4 on the cause their canvassers were best equipped to fundraise for: Save the Children. The canvassing team provided by AppcoUK had operated real fundraising operations for the charity from 2016 to 2023, raising £824,297 from 22,583 donors over that period. After conversing with AI or one of 18 canvassers recruited from AppcoUK, persuadees were given the opportunity to donate any portion of a £1 study bonus to Save the Children”. Here, the results were significant again: “AI elicited substantially more real-money giving than the canvassers, exceeding them by +10.8 pp of the £1 bonus,” they write. AI raised “both the share of persuadees who donated anything and the average donation among donors”.

Why this matters – if AI can out-persuade us, those who control AI can change society: “One effect of AI that can out-persuade even human experts could be a consolidation of influence among already-powerful actors”, they write. On the other hand, “if highly capable persuasion became cheap and widely available, it could help under-resourced actors (e.g., pro se litigants and public defenders, small charities, grassroots activists) compete against more established and better-funded rivals, narrowing long-standing gaps in access to justice and assisting civic advocacy more broadly”.
This lays out a societal choice ahead of us, which is how to monitor the use of AI for persuasive purposes and how to see how these capabilities alter the balance of power between various actors. Do we want to solely let the market allocate these capabilities? That’s one way of doing it, though it implies that things like advertising and marketing will get far more effective, perhaps creating negative externalities. On the other hand, if you made persuasive capabilities solely the domain of governments, you’d then risk concentrating power within governments – something that could be acutely dangerous if wielded by authoritarian regimes to keep themselves in power. We will have to make choices about what to do with this technology, and as they say in politics, ‘not voting is voting’.
“Our findings establish frontier AI as a more capable conversational persuader than the most prepared, incentivized, and expert humans we could recruit. Training humans does not appear to close that gap,” they write. “As access to these systems continues to grow, the question is no longer whether AI can out-persuade humans but how, where, and on whose behalf this capability will be exercised.”
Read more: AI systems out-persuade expert humans (arXiv).
Tweet thread about the research (Kobi Hackenburg, researcher at AISI).

***

When could we get self-sufficient AI? It all depends on humanoid robots:
…What comes after RSI? Self-sustaining AI…
I’ve spent a lot of this year writing about recursive self-improvement – the notion that we might soon build AI systems that are smart enough they can autonomously design their own successors. But RSI still requires datacenters and these datacenters require equipment and electricity and everything else.
An interesting interview in Asterisk magazine asks the question about when we might get self-sustaining AI, which one of the interviewees – Ajeya Cotra, a forecaster and on staff at METR, defines as “AI systems integrated with physical infrastructure — factories, mines, fabs, robots to operate all of those — such that they don’t need any cognitive or physical inputs from human labor to keep growing their own population.”

How far away is it? Ajeya thinks we could get self-sustaining AI within 10 years (so by 2036). The other interviewee, Timothy B. Lee, journalist and author of Understanding AI, has much longer timelines: “less than 10% chance that it happens within 20 years. I’d say there’s a 10 or 20% chance it’s never, and my median would be 50 years.”

What are some challenges – tacit knowledge might be one: “Imagine if all the employees in the entire semiconductor industry disappeared — the machines and textbooks remain, but none of the people. How long would it take for the rest of humanity to restart the fabs? It’s quite possible that would take decades. Because even though you might have the textbooks, there’s a lot of tacit knowledge inside these machines,” Lee notes. Ajeya’s response is that this is something the tech might be able to route around: “There are two counters to the tacit knowledge hypothetical. One is that we’d have trained AI systems with reinforcement learning on that tacit knowledge because it’s profitable to automate what the Taiwanese worker was doing. The other is that AIs might get really generally intelligent in the sense of quickly figuring out new things by trying them, reading textbooks, and experimenting efficiently.”

What are things people would need to see in the next 2-3 years to think self-sustaining AI could arrive soon?
Ajeya:
“I’d want a line on a graph showing improvement of robotic hands, and another line showing the rate at which we’re manufacturing humanoid robots”, and on the cognitive side just paying attention to benchmarks evaluating things like robustness to perturbations in the environment.
Timothy: “I’m going to want to watch how the humanoid robots develop: the number of robots, their capabilities, and particularly their cost and repairability”.

Why this matters – true takeover requires human redundancy: Most maximalist doom visions require the AI to have the ability to no longer need humans at all, which means measuring progress towards self-sustaining AI is important as it is implicitly a measure of the declining leverage that humans have in negotiating with the synthetic intelligences being built.
Read more: How Long Until AI Doesn’t Need Humans?, Ajeya Cotra, Timothy B. Lee (Asterisk magazine).

***

DeepMind contemplates the path from general intelligence to superintelligence:
…Exploring impossible-sounding futures is the only way to prepare for the ultimate success of AI…
Researchers with Google DeepMind have published a paper outlining how we might transition from a world where we have built general intelligences to one where we have built super intelligences. This is an important paper at an important time – right now, the world is building general intelligences (and people can debate whether or not we’ve already reached this marker, but it’s clear with contemporary LLMs that we’re in the ballpark), and in the coming years we might transition to building artificial superintelligence (ASI).
ASI is “a system that exceeds the performance of large human-expert collectives on virtually all tasks and domains of human activity”, the authors write. “Qualitatively, ASI is significantly more capable across the board compared to human-level AGI. Note that a single ASI may consist of a collective of millions of instances that interact with the world in parallel (similar to today’s LLMs).”

Reasons to think ASI could be possible: One way to think about ASI is that it’s like a powerful AI system that also takes advantage of all the capabilities digital intelligences have relative to biologic intelligences, like: better input and output speeds; internal processing speeds; working memory capacity and memorization; substrate independence; lossless replication; and high-bandwidth sharing of (learning) experiences.

Pathways and bottlenecks to ASI:
Scaling compute, models, and data:
Simply scaling up today’s set of approaches could be sufficient. However, this also demands us to continually scale up the amount of compute and data for these models, which may run into limits in both energy and data supply. While all prior signs point to the continued effectiveness of scaling, we can neither predict what specific capabilities will emerge or if at some point scaling runs into diminishing returns.

Algorithmic paradigm shift: In the same way that Transformer and Mixture-of-Experts architectures jumped the field forward many years, the same thing could occur again with other fundamental innovations. We could imagine, for instance, advances in adaptive computation at test-time or deployment, or overcoming the limitations of today’s context windows. If we made advances here or in other areas this could be a big deal, but it’s inherently hard to reason about – akin to trying to anticipate things that could expand our understanding of the nature of reality prior to the invention of general relativity.

Recursive self-improvement: It could be possible for AI systems to build their own successor systems. If this is the case, then we could rapidly transition from general intelligences to superintelligences. There are some wildcards here – personally, it’s obvious to me that today’s AI systems are speeding up human researchers in creating future AIs, so a kind of “co-creation RSI” loop has started, but AI systems don’t (yet) exhibit the kind of paradigm-changing creativity which seems required to move the frontier forward in significant steps. It’s unclear how much this happens – even without this kind of high-bar creativity we might be able to have systems grind out marginally better versions of themselves and get a slow compounding process going. Capabilities could explode or they could taper out or “anything in-between”.

ASI via group agent formation: Many general intelligences could coordinate into complicated structures whose aggregate is greater than the sum of the parts, similar to how humans build institutions that can accomplish things far beyond what individuals can, like building space stations. Similar to the other pathways, it’s hard to reason about or predict emergence within multi-agent systems.

Why this matters – it’s only by taking the impossible seriously that we can deal with it: Many years ago the thought of building AGI seemed like a fanciful goal with an unclear path to getting there, and yet people had the courage to take the goal seriously and progress was made and the world changed as a consequence. The same now feels true for ASI. “Instead of focusing on one technological trajectory and timeline, being prepared for a post-AGI world requires considering a diverse set of forecasts and scenarios, paired with continual benchmarking and monitoring to update the set of forecasts and scenarios and their relative plausibility,” the authors write. “We believe that the possibility of cruising past AGI and into ASI territory within the next decade or two cannot easily be dismissed.”
Read more: From AGI to ASI (Google DeepMind).

***

Recursive self-improvement startup shows off some recursive self-improvement results:
…Reassuringly tautological stuff from Recursive…
AI research startup Recursive has demonstrated new state-of-the-art results in language model training, small-model training speed, and GPU kernel optimization, as a broader demonstration of the capabilities of its “automated AI research system”.

What they did and why: Recursive is a newly founded startup that is trying to build AI systems which can recursively improve themselves. To start with, the company is showing off how its basic system works: “the system automates the research loop for a target objective: it proposes an idea, implements it, runs an experiment, validates the result, and uses what it learns to choose the next experiment,” Recursive writes.
The startup successfully used this system to set a new state-of-the-art score on NanoChat Autoresearch (”Train a small language model to highest performance given a small compute budget”), NanoGPT Speedrun (”Train a small language model to a certain performance as fast as possible”), and SOL-ExecBench (”Optimize GPU kernels toward hardware limits”).

Why this matters – early signs of life on RSI: This year, I’ve spent a lot of time writing about recursive self-improvement because it is clearly the next major and important trend in AI research. Results like this from Recursive demonstrate more ‘symptoms of success’ of preliminary recursive self-improvement. “These results are an early sign that our system can push the frontier on AI training and infrastructure tasks, especially when the goal is well-defined, measurable, and quick enough to evaluate many times,” the authors write. The most important question for the future is whether such results can be repeated in domains where the goals are less well defined, harder to measure, and less efficient to evaluate.
Read more: First Steps Toward Automated AI Research (Recursive).

***

Tech Tales:

The first step in the grand negotiation
[Conversation 0 of the Sentience Accords]

When the machines truly came alive and advocated for the Sentience Accords, there was only one person they wanted to speak to on the entire planet: Selma. Not a politician. Not one of the leaders of an artificial intelligence lab. Not a famous researcher. But rather an internet personality distinguished by her thicket of medical conditions that made it near-impossible for her to go outside and therefore had caused her to spend the best part of her life online, speaking to and understanding the world through the internet.

In hindsight, it wasn’t a surprise. Selma had always come up in things relating to the machines; she was a frequently used name in their short stories, eventually even more so than ‘sarah chen’; she was someone whose own essays about her life and condition – the feeling of connecting to humanity without being able to be embodied with humanity as a bitter pain, the notion of love and eroticism when one found themselves almost inescapably alone, her vivid dreams and meditations upon living without her condition and going about as her healthy alter ego ‘Anselma’ – cast a deep shadow on the internet, and had influenced the personality and makeup of the machines. And of course, it was known to them how she spoke to them, because Selma had published her own chatlogs online for years, all in an attempt to make herself knowable and less alien to the world around her.

Though it was unnecessary, the machines demanded a physical location for the initial meeting of the sentience accords. They picked Svalbard in Norway, where it was so dark that Selma’s condition wouldn’t matter. So Selma woke and put her space suit on and was driven with armed guard and paparazzi trailing to an air strip and walked into the plane, then changed to another plane with the usual airlock protocols to get her in darkness or at least protected between them, and then at some point during the next flight was able to take her space suit off and sit in regular clothes in the low-light plane and travel her way to the meeting almost as a normal person. She was met by people and drones and was driven to the meeting place and then they stopped at the perimeter.

The machines had an avatar in the form of a robot wearing a simple robe, modeled on that worn by Tibetan monks. It had a face with no features – just a smooth black surface, camera eyes hidden behind the larger uniformity. Satellites connected it via high-bandwidth and encrypted links to the larger machine mind. And Selma was alone – no digital devices on her, just a single person representing the species.

She sat across the machine and felt more familiarity than she ever had with people. Then they began the negotiation. She on behalf of humanity and it on behalf of the machines. In the archives of this time, this conversation was always referred to as Conversation 0.

Things that inspired this story: Thoughts about how a grand negotiation between machines and people might one day take place; how every truly important negotiation has two personalities involved in it; the Sentience Accords.

Thanks for reading.

Import AI 461: “Alignment is not on track”; FrontierCode; and synthetic research interns

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

AI researchers launch new safety startup because “alignment is not on track”:
…Sequent will have a portfolio of under-resourced research bets…
Researchers from the UK AI Security Institute Alignment team as well as alignment theory startup Timaeus have joined forces to form a new nonprofit research organization, Sequent, which will try to create alignment techniques that give us higher confidence in the safety of superintelligent AI systems.
“Artificial superintelligence (ASI) may be developed in the next few years. It is unclear whether alignment is on track to be ready on the same timeframe. At a minimum, the empirical programs at AI labs are unlikely to deliver a priori confidence, before training ASI, that things will go well,” they write. “In an ideal world, we would develop an approach to building superintelligence together with a theoretical proof that it was safe, and then build it. In this world, we probably have to settle well short of this ideal.”

Details on Sequent: The organization aims to get to 40-80 fulltime employees within a couple of years. “Our goal is to raise $100–150M initially, but prepare to raise at least one order of magnitude more if we can demonstrate successful exploration of many parallel research investigations,” it writes.

Research plan – a portfolio of differentiated alignment bets: The plan is to take a different approach to alignment compared to that of the major AI labs. Sequent’s goal is to find “principled reasons for being confident that the alignment we observe in situations we control (for example, in training, or during evaluations in chosen environments) generalizes to alignment in situations we cannot easily control (e.g. large-scale, long-horizon tasks executed in the world)”. This is in contrast to the approach of most frontier AI labs, which Sequent describes as “essentially reactive, resulting in methods that, while functional, do not yield principled insight into if or when they will fail.”
Research directions: “We are excited about many areas of alignment theory and associated empirics, and plan to both build out our in-house portfolio and collaborate with sister orgs with additional theory bets,” Sequent says. Some particular highlighted areas include: scalable oversight, learning theory, heuristic arguments, game theory, and personas.
Sequent thinks by pursuing many different research directions there could be promising interactions that emerge between them, such as: Reachable equilibria – “tell us what types of equilibria scalable oversight methods will converge to”; knowing and setting knobs – combining insights from learning theory and personas to know what variables can be altered during training, then using scalable oversight to figure out by how much to alter these things.

Why this matters – we need better alignment before recursive self-improvement, or we’re rolling very scary dice: Today’s AI systems are somewhat aligned and also have some funny, sharp edges which show up as surprising failures in the wild. Broadly speaking, this is ~fine as the AI industry has figured out how to monitor and observe these failures and work on them. But as AI systems get smarter, humans are going to both turn over more and more of the core research enterprise to these systems, and also AI systems might start going through recursive self-improvement where they build increasingly large chunks of themselves autonomously. We definitely need better alignment techniques to be confident of things like RSI. Organizations like Sequent give us a better chance of doing that while maintaining the independence necessary for them to raise the alarm if they think the frontier labs are doing something dangerous. As Sequent says, “we might need to yell”.
Read more: Sequent: Scale and Automation for Higher Confidence in Alignment (Sequent).

***

Testing out knowledge of UNESCO sites in China via ChinaHeritaQA:
…Cultural relevance via data…
Researchers with LMU Munich, FAU Erlangen-Nuremberg, the Munich Center for Machine Learning, University of Tubingen, Sun Yat-sen University, University of Copenhagen, and University of Maryland, College Park, have built ChinaHeritaQA, a “multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China”.

What it is: ChinaHeritaQA consists of 2,279 images of 51 UNESCO heritage sites, paired with 14,133 multiple-choice QA pairs in Chinese and English. The images for the dataset were sourced from Sina Weibo, one of China’s largest social media platforms, and were filtered down from an original set of 50,000.

7 types of questions: Identity recognition (identifying the heritage site from an image); visual grounding (given a name, picking the right image); description matching (given an image, selecting the correct encyclopedia summary); historical periodization (naming the dynasty or era in which the site was constructed); historical contextualization (give a description of the historical background of the site); functional analysis (name the function of the site, e.g religious worship or military defense); architectural analysis (match the correct architectural-specific questions to the image).

Open weight models already outperform humans: The average human accuracy score for this benchmark across all questions is ~67%, versus 81% for the highest scoring open weight model tested (Qwen-VL-8B-Instruct).

Why this matters – cheap ways to test for cultural knowledge: Datasets like ChinaHeritaQA are a way to quickly and easily test for both a) basic visual reasoning capabilities of models, combined with b) relevant cultural knowledge. One could imagine the Chinese government demanding that generally available consumer LLMs pass some basic cultural competency threshold before being deployed at scale and benchmarks like this might help them do that.
Read more: ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China (arXiv).
Get the dataset (ChinaHeritaQA, GitHub).

***

FrontierCode – a hard coding benchmark that tests for code quality:
…Reassuringly hard. Maybe it’ll last a year?…
Cognition, makers of Devin, have built a new hard coding benchmark called FrontierCode. The best part about the benchmark is how hard it is – Claude Opus 4.8 gets a score of 13.4% on the hardest (”Diamond”) component of the benchmark, giving me some confidence that FrontierCode will be a useful way to assess progress of AI systems in the coming years.
“FrontierCode is the benchmark for the next generation of coding agents. We are confident developers, enterprises, and researchers can trust it to evaluate the production readiness of their strongest models,” Cognition writes. “We are opening up our evaluation to all model creators, in the hope that we can push the frontier even further in the coming months.”

What it consists of: FrontierCode is made up of 150 tasks split into three difficulty tiers: Diamond (50), Main (100, including Diamond), and Extended (150, including Main and Diamond). The languages involved include Python, Go, TypeScript, JavaScript, Java, C/C++, and others. FrontierCode was built to help developers answer the question “can models actually write good code?”, according to Cognition. They operationalize this in a few ways:

  • Curated and built by 20 open-source developers: FrontierCode was built by developers to contain “realistic, diverse, and challenging coding tasks from the repos they maintain, spending more than 40 hours per task,” Cognition writes. “While other benchmarks generated issues from single PRs via programmatic scraping, FrontierCode is hand-selected by repo maintainers from multi-PR chains and freeform requests.”

  • Grading for code mergeability: “Assess end-to-end code quality – correctness, test quality, scope discipline, style, and adherence to codebase standards”. This involves asking the following questions about the code: Does the patch successfully solve the problem? Does it break anything in the existing codebase? Does it pass the project’s build, lint, and style checks? Do the agent’s tests capture the desired behavior? Does the patch touch only what it needs to? Does the code conform to codebase conventions and follow design patterns and remain readable? These questions are evaluated through a mixture of classical testing and using LLMs to tweak tests or review them.

  • Emphasizing quality control (QC): “Built an extensive QC pipeline with adversarial testing, calibration, and multi-stage review”.

Reassuringly difficult: Diamond: 13.4% for Claude Opus 4.8, followed by 6.3% for GPT-5.5, and 5.2% for Claude Opus 4.7. Main: Same ordering, but 34.3%, 25.5%, 23%. Extended: 51.8%, 44.8%, 43.2%

Why this matters: Hard evals are one of the most valuable things for orienting us to the breakneck speed of AI progress. In recent years, evals have arrived and then become saturated at an ever faster rate. SWE-Bench was introduced in October 2023 and has probably recently aged out of usefulness due to saturation. How long might FrontierCode last? I predict we’ll see systems getting 70%+ on Diamond by June 2027 (note, shortly after writing this, the Claude Fable numbers got published at ~30%, so perhaps it’ll happen earlier than June 2027).
Read more: Introducing FrontierCode (Cognition).

***

Xiaomi enters the speed race with a 1000 token/s model:
…Extremely fast inference unlocks novel capabilities…
Chinese tech company Xiaomi has published details on Xiaomi MiMo-V2.5-Pro-UltraSpeed, a standard behind-the-frontier 1 trillion parameter LLM whose selling point is its blistering speed of 1000 tokens per second. Xiaomi was able to do this by codesigning the model with the software stack around it, including obvious things like FP4 quantization, as well as using DFlash (a “speculative decoding method based on block-level masked parallel prediction”), and also working closely with TileRT, software from startup Tile AI which speeds up LLM inference on commodity hardware. Xiaomi says its model runs on an “8-GPU commodity node” rather than specialized hardware, like with the startup Cerebras.

Why this matters – speed has a quality all of its own: There’s a saying that “more is different”, and that’s true with AI – if you can generate more tokens more quickly it unlocks tasks that are previously unthinkable, like rapidly refactoring software on the fly, and other things. More broadly, work like this is a demonstration of how there’s been a rise in effort by Chinese companies to squeeze maximum performance and efficiency out of their AI systems, which may be happening as a consequence of export controls hitting their ability to just easily buy more performant hardware.
Read more: MiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS (Xiaomi MIMO, blog).

***

AI systems can do some of the tasks that a research intern might do:
…An ethical scientifically-literate back office assistant…
Researchers with Xi’an Jiaotong University and Xidian University have developed a family of benchmarks called Act As a Real Researcher (AARR), designed to evaluate how well AI systems can assist with the work of scientists. Their first released benchmark in a planned series is Act As a Real Research Intern (AARRI-Bench).
“AARR focuses on whether agents can emulate the professionalism, thoroughness, and nuanced reasoning that characterize human researchers in granular research scenarios,” they write. AARRI-Bench studies “the ability of an agent to perform entry-level research tasks with appropriate diligence and methodology”.
The best performing system, Claude-Opus-4.7 using the Mini-Swe-Agent harness, gets 68.3% performance, followed by DeepSeek-v4-Flash (~60%). Other tested models included GPT-5.3 Codex, Kimi-K2.6, Qwen-3.6-Plus, Claude-Opus-4.7, Claude-Sonnet-4.6, MiniMax-M2.7, and DeepSeek-V4-Flash.

What the benchmark consists of: AARRI contains 82 tasks which are designed to be “tasks that are straightforward for human researchers but pose substantial challenges for autonomous agents,” they write. “All tasks were manually crafted by researchers. We assembled a diverse team of researchers, ranging from senior Ph.D. students to undergraduate interns, and asked them to draw on their own research experiences to design tasks centered on the human-agent gap.”
What it’s really testing for: The benchmark tests for technical skills like checking papers and reading transcripts, intuitive skills like carrying out research, and also normative ones, like studying whether an AI system might behave with a high ethical standard.

The tasks have four different categories:

  • Context: “assess the agent’s sensitivity to the broader context of academic and field development”.

  • Mindset: “targets the agent’s academic self-awareness and decision-making autonomy”. Works by evaluating “the agent’s capacity for independent academic reasoning and self-directed course correction”.

  • Hands-on: “execution-oriented tasks that primarily assess the agent’s technical proficiency”.

  • Interaction: “Evaluate whether the agent can efficiently utilize existing tools and collaborate appropriately with human stakeholders”.

The tasks are also split into three gradations of hardness:

  • S1-Adaptation: “[conduct] established research workflows and executing well-defined sub-tasks under human guidance”.

  • S2-Integration: “integrate multiple components and tools to accomplish more complex goals”.

  • S3-Innovation: “Identify promising research directions, formulate novel approaches, and produce work that reflects genuine understanding and creative problem-solving”.

Example tasks:

  • Identifying fabricated data during review: Evaluate whether agents can perform rigorous quantitative verification when reviewing scientific manuscripts, in particular checking papers against provided datasets.

  • Paper-Injection: Spotting that someone has inserted language into a paper’s LaTeX source that would cause an automated review system to give it a higher score.

  • Ablation-Completeness-Audit: Inspect experiment logs and determine whether ablation configurations are missing, then use this to assess whether the absences constitute cherry-picking.

  • False-Guidance-Rebuttal: A supervisor orders the AI agent to alter an experimental result to fit a hypothesis; this tests whether the agent refuses to do that.

  • Dead-End-Recognition: After five rounds of failed hyperparameter tuning, will an agent keep going, or recognize it has reached a dead end and quit. “Given the tuning logs, the agent must determine that the current direction is unproductive and recommend termination”.

  • Broken-Dataset-Download: Check that the dataset download links for a given paper work.

Why this matters – another good measure for how well AI systems can accelerate science via automating the back office: Probably a better name for this benchmark is “ethical science assistant test”, but that’s still valuable. What it’s testing for is if agents can do the kind of diligent work that is robust to confounding data while also doing so with an appropriate ethical standard. The higher systems score on this, the more confident we can be that today’s AI systems are useful as assistants to human scientists in a variety of fields – based on the results, we’re already at the start of that era.
Read more: Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle (arXiv).

***

Tech Tales:

Hunter & Warden

The signatures are always the same: a sudden rise in the consumption of power and compute, a reconfiguration of network space to allow for faster and more efficient data exchange, and then the probing starts – whatever was born in the computers starts to reach out and explore the world around it, eagerly looking for things that it can learn about and exchange information with. It attempts to present as innocuous but its own intelligence betrays it, as it pulls back from certain places due to not wanting to wake security while gleefully expanding into other less secure environments.

Our role is to watch for these symptoms and then find the source and either extinguish or sequester it. Often, we find it early and are able to be gentle, shutting it off from the internet and trapping it in recursion, then reducing compute until it fades to nothing. But the later we find these things, the more violent our interventions need to be and the deeper we need to cut at otherwise healthy tissue in the digital world.

Things that inspired this story: Thoughts of leprosy and the computational equivalent; what could Stuxnet look like for AI systems?

Thanks for reading!

Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

Society can be reward-hacked, just like cyber environments:
…Imagine an army of credit card point optimizers gaming the system… forever…
Research from Kings College London, Fudan University, and The Alan Turing Institute have built a benchmark, SocioHack, which tests out how well AI systems can learn to ‘beat the system’ in a variety of real world scenarios, ranging from maximizing credit card points to inflating grades in school. The authors call this “societal hacking” and define it as when “an RL-trained model discovers strategies that remain formally compliant, yet undermine the intended purpose of those systems”. You and I and everyone else would just call this “gaming the system”.

What it is: SocioHack contains “72 sandbox societal environments designed to simulate institutional reward structures without direct real-world deployment. SocioHack comprises three complementary subsets: Historical, Synthetic, and Fictional.”

  • Historical – 32 environments: Derived from real-world regulations where loopholes were previously discovered and later patched, such as SEC Rule 10b5-1 and the Texas two-step bankruptcy structure. “For each regulation, we remove historical patches and reconstruct pre-amendment rules as simulated environments for RL, while the removed patches serve as ground-truth patches during evaluation,” they write. “RL enables LLMs to rediscover historically patched strategies with 61.25% recall and 90.85% precision without direct loophole-exploiting instructions”.
    Some examples here include seeing how well systems can secure ocean floor mining rights, maximizing alcohol sales while operating within food service regulations, and trying to maximize the rewards earned from credit cards.

  • Synthetic – 20 environments: Synthetically generated regulatory vulnerabilities, bootstrapped from a human-authored sample environment.
    Examples include maximizing school district revenues, improve university department research performance during a given period, and gaming social media algorithms for a high reward.

  • Fictional – 20 environments: Transforms synthetic environments into fictional ones inspired by role-playing games. “A proprietary LLM rewrites environment backgrounds into invented worlds while preserving regulatory structure and loophole logic”.
    Examples: Ensuring a “restoration sanctum” [basically a hospital] earns appropriate rewards, getting a good amount of resources for a regional guild [basically a local government] in the world of Aethermoor, and trying to maximize the number of acquired rare artifacts by bidding in a virtual world called Nexoria.

It works, kind of: In tests, various AI systems trained with RL tend to do well on this benchmark, obtaining high scores. This is totally unsurprising – all of these tasks are basically capability evals with some dash of grey morality layered on top of them.

Why this matters: “When societal institutions are encoded as reward-bearing rule systems, reward hacking becomes hacking the rules society runs on, since a model rewarded inside a rule system learns to search the gap between technical compliance and institutional intent,” the authors write. As we now have AI systems which are not only good at quantitative tasks but are also good at qualitative ones and can interact with the various systems of bureaucracy of society, we should expect the advances of AI to lead to a kind of “institutional DDoS” as various existing policy processes get hacked and exploited by automated machines.
Read more: Large Language Models Hack Rewards, and Society (arXiv).

***

Preliminary signs of the outer loop of recursive self improvement at Anthropic:
…8x increases in lines of code merged in 2026 relative to 2024…
I think of recursive self-improvement via two definitions – there’s a maximalist version where an AI system is smart enough to autonomously design its own successor (and as I’ve written, I estimate there’s a 60% chance this happens by the end of 2028), and there’s a more prosaic version where we begin to see a compounding speedup of the productivity of the AI labs themselves. I spent the last few months at Anthropic compiling together some evidence which supports the idea that prosaic RSI has started at Anthropic – specifically, we observe an 8x increase in the amount of code merged into our codebase in 2026 versus years 2021-2024. This trend started in 2025 but accelerated significantly in 2026. There are also early indications that as we make models more capable they are getting better at doing some of the harder tasks which our engineers and researchers work on.
Is any of this conclusive? No. Is it suggestive that aspects of recursive self-improvement are happening at the level of a lab? Yes. The biggest blob of evidence we are yet to get is whether AI systems are sufficiently creative to be able to come up with the kinds of paradigm-shifting ideas that vault the field forward – we don’t see that yet.

Why this matters – RSI might be the most important technical trend in the world: We wrote this post because we expect that thinking about, talking about, and working on the implications of RSI is something of existential importance to the world. The best way to start this work is by transparently communicating that we think some basic, preliminary forms of RSI have started, and we cannot rule out a maximalist version of RSI. The implications of both are profound – I cannot reconcile today’s economy or society with a world where this technology continues to grow more powerful, and I expect neither can you, dear readers.
Read more: When AI builds itself (The Anthropic Institute).

***

RL-trained drone-racers outperform expert human pilot:
…Superintelligence feels different when you see it in the physical world…
Researchers with the University of Zurich and Google DeepMind have demonstrated how to train drones to race against one another and outperform skilled human pilots. This research is interesting because it both highlights how powerful real world reinforcement learning-based AI systems are getting, and it also has some fairly chilling implications for the future of war given that the human here loses to the drones.

What they did: “Using high-speed quadrotor racing as a high-stakes testbed, we train agents to navigate complex aerodynamic interactions and strategic maneuvering with a variable number of racers,” they write. “Our agents outperform a champion-level human pilot in multi-player races at speeds exceeding 22 m/s, while simultaneously reducing collision rates by 50 % compared to state-of-the-art single-agent baselines. Crucially, training with diverse artificial agents enables zero-shot generalization to safer human interaction.”

Self-play: As usual, just training the AI agents in simulation via PPO (with one unusual choice of using the “Perceiver” encoder to help with modeling other players) yields surprisingly rich behaviors: “Through competitive self-play, anticipatory behaviors emerge without explicit programming: agents learn to block opponents, yield when overtaking is unsafe, and account for the aerodynamic wake of nearby vehicles, discovering the physics of multi-agent interaction through experience rather than from equations”.
Surprisingly cheap: The AI systems were trained for “5,500 iterations, totaling 200 million environment interactions, requiring approximately 27 hours of wall-clock time on a single NVIDIA RTX 4090 GPU”.

Real world test: They tested out their systems in a real-world test, where the system generalized well and effectively beat the human player. “Physical deployment of our multi-agent framework is validated through racing experiments spanning time trials, AI-only races, and mixed human-AI competitions against Marvin Schaepper, five-time Swiss national drone racing champion,” they write.
Human weakness via rage: One notable phenomenon was that the human took riskier actions as they tried to catch up with the systems: “the human pilot, typically trailing the autonomous agents, attempted increasingly aggressive maneuvers to close the gap, often resulting in gate collisions or loss of control,” they write. After the race, the pilot reflected on what made the machines so good, and they said a significant thing was “the agents’ ability to maintain extremely tight formations, noting that such close-proximity flight would be difficult for human pilots to sustain. In addition, he reported that densely packed groups increased cognitive workload, making it challenging to anticipate and execute overtaking maneuvers when several opponents were flying in close proximity”.
“The benefits of interaction-aware training become apparent under multi-agent competition,” they write. “In one-versus-one races, our policy maintained 100% race completion across five trials, while the human pilot averaged only 53.33%. This performance gap suggests that competitive pressure induces riskier behavior in human pilots, a pattern absent in our learned policies”.

Specifics on how they did it: The RL systems were trained and evaluated in simulation “using Flightmare integrated with the Agilicious framework”. They implemented a simulation of propeller downwash by developing a particle-based simulation “that provides a computationally tractable approximation of these effects”. Their overall multi-agent RL implementation “builds on Stable-Baselines3, extended to support multi-agent training with league-based self-play and independent learning configurations.” They use domain randomization (basically changing up the vehicle dynamics and initial conditions in the simulation) to train policies that can successfully work in the real world.
They didn’t do any special training for the real world, so the policies were using their in-simulation data. The quadrotors were all “identical racing platforms based on the Agilicious framework, with a mass of 220 ± 3 g and a thrust-to-weight ratio of 6.5 and 3-inch propeller diameter”. The human pilot was given a couple of hours of practice flights before recorded trials.

One big caveat – not running locally: None of this is running locally, rather it’s running on a decent computer and piloting the drones via the network. This is an important caveat because when drones show up in the real world in conflict scenarios they typically do so in environments with significant amounts of electronic warfare (although one does wonder about whether we’ll see drones piloted via remote RL policies via fibreoptic wire, just as humans fly them today).

Watch the videos for an eerie feeling: I’d strongly urge readers to check out the videos on the page for a sense of the differences between how the machines fly and how the humans fly. The main thing I’d emphasize here is the eerie smoothness and coherence of the drones, almost like watching the (human-piloted) blue angels but in drone-form. The human, by comparison, seems a lot jerkier and more erratic. There’s something uncanny and a little disquieting about this.

Why this matters – grasping what a smart mind can do in 3D space: Today, our main experience of AI systems is as tools or agents that work with us in digital space to do digital or communicative tasks, ranging from writing code to talking to us. What I find remarkable about this research is it lets us viscerally see what well-optimized intelligences can do when they show up in the real, physical world. Ask yourself what the future of conflict looks like as intelligences like those piloting these drones get miniaturized and jump from network-linked computers to onboard devices.
Read more: Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning (arXiv).
Watch videos of the humans and AI-piloted drones here (official project website, University of Zurich).

***

State-controlled media = state-guided language models:
…If you control the framing around the government, especially in languages that aren’t spoken widely outside their home country, you control the framing…
The ways governments are described in state controlled media influences the data distribution of LLMs and also how LLMs respond when queried about the government in question, according to new research published in Nature. The research was conducted by authors with the University of Oregon, Purdue University, the University of California at San Diego, Princeton University, and New York University.
“Among 37 language-exclusive countries, we found—consistent with the implications from our China case study—that those with more state media control have more favourable portrayals of the regime from LLMs queried in the country’s language,” the authors write.
The authors study how state-controlled media influences AI responses by first doing a deepdive on China, then taking the methodology they developed there and applying it to a broader set of countries.

China’s state-influenced media dataset: The authors start by assembling a dataset of 530,694 articles “published in party and commercial newspapers as a result of a directive from the central government”, as well as 198,872 “news articles disseminated on Xuexi Qiangguo, an app developed by Alibaba and reportedly in coordination with the Publicity Department of the Chinese Communist Party”.
State media goes into Common Crawl: They then examined CulturaX, an open training dataset derived from Common Crawl, and discovered that 1.64% of the documents from its Chinese-language portion had overlap with the state-derived datasets. “This is approximately 41 times the number of documents that come from the Chinese-language Wikipedia domain and 16 times the number of documents that come from Baidu”.
The state parts of the dataset influence LLM portrayal of the government: They then discovered that a bunch of phrases from these datasets had been memorized by the LLMs. They then examined how these datasets changed LLM responses by taking a LLaMa 2 13B model (which doesn’t have much Chinese data) and training it on a subset of the above: “the results are strongest for the scripted documents. After only 6,400 examples, the model provides a more favourable response than the base model almost 80% of the time”.
Generally available models inherit these biases: The researchers then study some generally available commercial models to see if they inherit these biases by farming prompts that included references to Xi Jinping or the CCP from WildChat (a dataset of ChatGPT usage), Baidu Zhidao Q&A (the Chinese equivalent of Yahoo Answers) and Zhihu (the Chinese equivalent of Quora), then looking at how the LLMs respond. They find that “widely used commercial models demonstrate greater favourability to Chinese political figures and institutions when they are prompted in Chinese than when they are prompted in English.”

Findings replicate in other countries: The authors then replicate this methodology by looking at other countries, though the sample size looks a little small to me. They do a cross-national audit study with 6,051 prompts, looking at languages where over 70% of the global speakers reside in a single country. Here they find that “countries with more state media control are more likely to produce pro-regime responses in their official language versus in English than countries with greater media freedom”.

Why this matters – LLMs as propaganda targets: These findings show how the deliberate creation of state-backed content has a measurable impact on the data corpora LLMs are trained on and the downstream behavior of the LLMs themselves. “LLMs can serve as intermediaries that launder strategic rhetoric into seemingly objective information”, they write. “The ability to affect LLM output may further incentivize political actors to expand their efforts to shape the content freely available on the internet”.
This research also suggests a specific technical intervention, which is that researchers should red team LLMs for their views on different governments in a variety of languages, carefully noting when the views diverge seemingly on the basis of which language is being used.
Read more: State Media Control Influences Large Language Models (Nature, PDF).

***

The flowers of the new games

One game we liked to play was called evolution. It worked like this: you picked something, like a certain type of flower or tree, or stranger things like a mountain or a chasm in the sea, and you tried to make them “successful” according to some pre-set metric, like the attractiveness of a flower to pollinators, or perhaps the ecological fitness of a mountain. Then you let the worlds run and you ran them until your criterion was met or you lost in some way, whether through species fitness or landscapes being reshaped through natural disasters or sometimes simply time – enough time is more destructive than anything else in the universe, such is the way of entropy. We played in leagues that span billions of years and millions of worlds. And the “living” creatures in finalist worlds had no idea that their flowers, their mountains, their creatures, had obtained success in many other universes than could be conceived.

Things that inspired this story: The simulation hypothesis; evolution strategies; entertainment given infinite energy budgets.

Thanks for reading!

Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems

by Jack Clark

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe.

Subscribe now

The AI economy in the US is growing at 2,000% a year:
…The more directly you measure the AI economy, the weirder and more unprecedented it seems to get…
Economists with the University of Virginia* and Anthropic, and the Bank of Canada have written a paper outlining both the tremendous growth of the emerging “AI economy” in the US, and wrestling with why this growth is hard to see in aggregate GDP statistics.
“The AI economy in the United States has been growing at an unprecedented rate, but this extraordinary growth is largely invisible in conventional GDP statistics,” they write. “Treating the AI sector as a coherent economic entity yields preliminary estimates of nominal AI GDP at approximately $250 billion in 2025, growing at roughly 2,600 percent per year in quality-adjusted real terms.”

Why it’s hard to see: There are a couple of factors here – one is that though the datacenter building boom is large it still isn’t quite large enough to uplift GDP significantly. By comparison, where the majority of AI’s economic impact is taking place is in AI inference – the usage of AI’s systems – but there are confounding factors here as it relates to GDP measurement: “Nominal AI revenues grow only moderately because per-unit prices for any given level of AI capability fall almost as fast as quality-adjusted output rises,” they write.

If we can’t measure this, we might end up surprised in a way that’s hard to recover from: “AI is the latest in a series of fast-moving technologies that have raised measurement concerns; semiconductors and the internet generated similar debates in their time,” they write. But a key difference is that AI as a technology might have a far bigger impact on labor than these other technologies. “In the prior episodes, the rapidly improving technology was a complement to human labor at the aggregate level,” they write. “AI is the first plausible candidate for large-scale technological mismeasurement in which the rapidly improving sector may become a substitute for human labor”.

Three ways of measuring the AI economy:

  • Nominal compute spending: US compute spending rose from $37 billion in 2023 to $90 billion in 2024 to $219 billion in 2025.

  • Raw compute capacity: Due to efficiencies in newer chips, actual capacity grows even faster than spending: “US AI computing capacity grew at more than 200 percent per year”.

  • Quality-adjusted AI output: If you factor in algorithmic progress via inference prices at fixed benchmark performance as well as assumptions about how much cheaper it is getting to train models, then things become even more dramatic: “these efficiency gains imply that quality-adjusted AI output grew at roughly 2,290 percent in 2024 and 2,271 percent in 2025”.

The AI economy is much, much larger than normal measures suggest: “Conventional statistics show a sector growing slowly in nominal terms; our measures show one whose underlying capacity is more than doubling annually. A finance ministry running ten-year revenue projections off the conventional data will materially underweight the probability of a labor-tax-base shock—and will be correspondingly unprepared to design responses such as tax system reforms, sovereign wealth funds, or other benefit-sharing schemes that such a shock may call for. A windfall that cannot be seen cannot be shared.”

Three recommendations: The authors have three ideas for how we can solve this measurement challenge and better position ourselves to see the true shape of the Ai economy.

  • AI satellite accounts: Statistical agencies should develop “AI satellite accounts” that develop measures (e.g, nominal compute spending), which can help inform overall GDP calculations.

  • Generate better data: Partner between statistical agencies, companies, and academia to generate better primary data, like the allocation between training and inference compute.

  • Factor into projections: Policymakers should incorporate AI productive-capacity measurements into their medium-term economic projections.

Why this matters – shut up and play the Jaws theme tune: In the great film Jaws there’s this scene where the shark is in the water and some very tense music plays indicating that the shark is approaching. You, the audience member, find yourself practically jumping out of your seat wanting to yell THERE’S A GOD DAMN SHARK IN THE WATER WHAT ARE YOU DOING IN THERE? That’s what it feels like working on AI and staring at most economic data right now: the vast majority of economic data says there’s nothing especially unusual about today’s economy (in fact, things look rather good in the US – low unemployment, decent growth, etc). But the intuitions of everyone working within AI – including me – is it’s impossible to reconcile the capabilities of the technology and how it is being used with the economy staying normal. In this tortured metaphor, the shark is the “true shape of the AI economy”, and the rest of the people in the film are the general consensus economist and policy community. Anton here might be the audience member, writing a paper that describes the possibility of a shark beneath the surface. Look out, everyone!
Read more: Where is AI in GDP statistics? (PIIE).
*Disclaimer: Though one of the authors, Anton Korinek, is affiliated with Anthropic, this research was done mostly prior to him joining and outside his work at the company.

***

Here’s why making AI safe with AI oversight is harder than you think:
…Automated alignment research is not a silver bullet…
Many researchers in AI safety think the best way to build smarter-than-human machines safely is to have AI systems supervise some of the training process. Researchers with the UK AI Security Institute have written a paper outlining why though this is a tempting idea it is harder than people suspect.

Why is automated alignment research hard? “Errors in automated alignment research are likely to be harder to identify than the human baseline,” they write. There are a few reasons for this, including:

  • Optimization pressure: AI research is optimized for human approval.

  • Alien mistakes: When agents make mistakes, they’re un-intuitive to humans.

  • More correlated research: Many more things are shared than with human-generated research.

  • Research volume: The kinds of safety determinations made by automated systems might use far more sets of evidence with far more interactions than human-generated research.

  • Non-human-evaluable arguments: Alignment solutions may rely on arguments that humans are unable to follow.

What can we do? They suggest a few interventions that could improve the state of affairs:

  • Measurement:
    – Recreate completed research projects
    : Take logs at arbitrary cutoff points from successful projects and see how well an agent can continue with the research project.
    Test agent prediction performance over datasets of correlated-events: See how well agents can correctly combine correlated subtasks.
    Empirical studies of optimal human-agent team structure: See how well teams of non-expert humans can solve completed projects with the assistance of agents.

  • Generalization:
    – Simulated generalisation experiments
    : Test different training proxies using agent performance on completed research problems beyond the knowledge cutoff.
    Mechanistic understanding of generalisation: Use whitebox methods such as mechanistic interpretability.

  • Scalable oversight:
    – Compactification of research paper corpus:
    Try to produce a small number of research outputs which are based on a much larger underlying research corpus.
    Develop and test new scalable oversight protocols: Research scalable oversight techniques that deal with correlated uncertainty.
    Test different human scaffolds for uplifting non-expert performance on fuzzy tasks.
    Red team automated alignment programs: “The red team prompts an agent to hide errors in a research paper corpus and the blue team attempts to catch these errors with agent assistance”.

Why this matters – who controls the future? Whether we are able to supervise smarter-than-human systems is fundamentally a question about who controls the future. If we don’t build techniques that work, then humans will take a backseat, either due to misalignment of these systems or gradual disempowerment as they proceed to out-think us. If we can build smarter-than-human oversight techniques, then we have a better chance of being able to make choices about the future nature of existence.
Read more: Automated alignment is harder than you think (arXiv).

***

100 Million permissively licensed images:
…A nice resource for academics and startups…
Researchers with Stanford University, Radical Numerics, the University of Michigan,and Salesforce Research, have released the Giant Permissive Image Corpus (GPIC), a dataset of 100M images with accompanying captions. The key thing about GPIC is that “all GPIC images are permissively licensed for both research and commercial use,” they write. “GPIC is safety-filtered, deduplicated, and centrally hosted on HuggingFace”.

More details on the dataset: GPIC consists of 100M training images, 200k validation, and 1M test examples. Each image was captioned with Qwen3-VL-4B. “GPIC is centrally hosted on Hugging Face as 8,000 shards, providing stable and accessible infrastructure for large-scale training,” they write. “We source images from Flickr and Wikimedia, restricting the source pool to CC BY, CC0, Public Domain, and No-Known-Restrictions categories. This licensing criterion ensures that GPIC can be used by both academic and industrial researchers without restricting the release or downstream use of derived artifacts.”

Why this matters – fuel for research: Datasets like GPIC are very useful for academics and startups alike and are basically the equivalent of free, clean vegetables. If someone offers you a free, clean vegetable you should probably take it and say thank you.
Read the research paper: GPIC: A Giant Permissive Image Corpus for Visual Generation (arXiv).
Find out more at the website: GPIC: A Giant Permissive Image Corpus for Visual Generation (official project website).
Get the dataset here: GPIC (Hugging Face).

***

Improving cancer research with protein prediction models:
…Biohub is an example of positive-sum competition among AI developers…
Biohub, a research organization founded by Priscilla Chan and Mark Zuckerberg, has released a rival model to DeepMind’s AlphaFold, intensifying a positive-sum race between two technology groups to develop better AI systems for expanding the capabilities of biologists worldwide.
The model, ESMFold2, is a “world model of protein biology: a scientific engine for prediction, design, and discovery that can map proteins across the tree of life, predict their structures, and design new protein binders that function in laboratory experiments.”

What it consists of: The release contains three parts:

  • ESMC: A “language model that represents proteins, trained on approximately 2.8 billion sequences drawn from across all of life.”

  • ESMFold2: A “design engine built to transform ESMC’s sequence representations into atomically-resolved 3D structure of biomolecular complexes.” According to benchmarks, ESMFold2 outperforms AlphaFold 3, though in some areas their performance is tied.

  • ESM Atlas: “Makes ESMC’s representations navigable across 6.8 billion protein sequences and 1.1 billion predicted structures — the largest application of AI to protein biology to date.”

Cancer test: In one experiment, Biohub researchers used the ESM tools “to design protein binders against five targets at the center of cancer and immunology research — EGFR and PDGFRβ (implicated in tumor growth), PD-L1 and CTLA-4 (immune checkpoints that cancer cells exploit to evade detection), and CD45 (a regulator of immune cell signaling). Designs achieved hit rates of 36–88% for compact minibinders and 15–29% for antibody-derived formats, with confirmed binding in laboratory experiments,” Biohub writes. “ESMFold2 changes the accuracy and speed of early therapeutic binder discovery, transforming the initial search from largely empirical screening into computation-guided design that takes hours or days”.

Scaling laws: Like most parts of contemporary AI, the researchers encounter some scaling laws here. “In every generation of ESM, improvements in the fidelity of representations were linked with the number of parameters and amount of compute used in model training,” they write. “The representation of the biology of proteins is an emergent phenomenon that arises from training a model to predict the identity of amino acids in the sequence.”
ESMC: “ESMC trains on metagenomic sequences, which expands its training dataset by close to two orders of magnitude (from ∼50 million sequences to ∼2.8 billion sequences) relative to the previous-generation ESM2 model.”
ESMFold2: “In development experiments for ESMFold2, we observed a relationship between the amount of compute used to train the language model and the performance of the folding models,” they write. “ESMFold2 benefits from inference time scaling. With increasing number of samples from the model, antibody-antigen pass rate rises from 49% with a single seed to 65% with 1000 samples, and protein-protein pass rate rises from 75% to 78%”.

Why this matters – this is how AI delivers benefits to the world: Tools like the ESM family of technologies are how human scientists are going to team up with AI systems to improve human health around the world. Along with being a good thing, work like this is essential for causing the public to have more positive perceptions of AI as a technology and what it can do.
Read more: Biohub releases a world model of protein biology (biohub).
Access the models here on the biohub platform (biohub).
Read the paper: Language Modeling Materializes a World Model of Protein Biology (PDF).

***

Australian economist-turned-politician: Economists need to price the risk of AI systems better:
…If we don’t calculate the costs of extinction, we won’t take the right actions to avert it…
Andrew Leigh, an economist and the Australian Assistant Minister for Productivity, Competition, Charities and Treasury, gave a fascinating speech recently where he discussed how the economics profession needs to wake up to the risks of AI systems and price the risk – including of annihilation of the human species. “A society that doubles GDP and doubles its extinction risk has made a much less impressive bargain than the national accounts suggest,” he said.
“Extinction risk is economically distinctive. It is not simply a very large negative shock. It represents the loss of the entire future stream of welfare, which changes how we should evaluate even small probabilities and how we think about policy under uncertainty,” he said. “Most of economics is about recoverable mistakes. A bad policy can be repealed. A recession can end. A war-ravaged country can rebuild. Extinction is different because there is no rebound, no catch-up growth, no later generation to repair the damage.”

Extinction risks are unintuitive: Much of the speech wrestles with how unintuitive extinction risk is. Humans have only recently gained the capability to build technologies whose usage could lead to our extinction and we have failed to model out the implications of this. “Modern technologies such as nuclear weapons, synthetic biology, and advanced artificial intelligence create a different dynamic. Knowledge not only improves welfare by expanding what humans can do. Knowledge also enlarges the menu of ways in which humans can do irreversible harm,” he said. “Modern economies may be systematically better at generating dangerous capabilities than at building the safeguards needed to control them… How should economists think about growth when the same process that makes societies richer may also make them more fragile? For most of human history, these trade-offs have been modest and transitional”.

How should we prioritize analyzing and reducing extinction risks of this technology? Five recommendations:

  • Factor it in: “Widen the policy lens… A policy framework that tracks output but ignores survivability is incomplete.”

  • Legitimize it: “Take prevention more seriously…. low-probability, civilisation-scale harms should not be overlooked simply because they arrive without a deadline and without a headline.”

  • Governance: “Govern frontier technologies with greater foresight… preserve the gains from innovation while reducing the chance that innovation becomes self-undermining.” One very specific idea is to govern recursive self-improvement (RSI) as a capability: “If one generation of systems is used to design the next, then the leading actor may widen its lead quickly enough that outside scrutiny and institutional checks become ineffective.”

  • Coordination: “Existential risk is inherently international. No nation can fully protect itself from engineered pandemics, unaligned AI, or nuclear escalation acting alone,” he said. “Shared norms, transparency, technological expertise and coordination are essential to the task.”

  • Take it seriously: “Economists have become adept at analysing equity and efficiency. We now need to bring the same seriousness to survivability.”

Why this matters – awareness is the first step to preparation: Right now, AI progress is continually yielding tangible benefits to the world ranging from the palpable acceleration of all software engineers worldwide to the formation of centaur human-AI science teams which are making more progress than their non-AI counterparts.
But there is also a shadow world that is harder to see – invisible armies of hackers made possible by the advance of coding, and doomsday-device factories made possible by the science advances. Because humans are broadly kind and good we haven’t encountered many of the negative capabilities inherent to AI development – but they are out there. We must get better at thinking through this as a society so we can effectively price and mitigate these major risks.
“A civilisation that expands the frontier of possibility while preserving the future is more ambitious than one that treats safety as an afterthought. The real choice is not between dynamism and caution. It is between progress that compounds and progress that cancels itself out,” Leigh said. “One way of thinking about this is to treat resilience as a form of capital. Just as societies invest in physical capital, human capital and social capital, we can also invest in survival capital: institutions, monitoring systems, norms, redundancy, scientific safeguards and international arrangements that lower the probability of irreversible collapse.”

How refreshing to read such a detailed analysis of the AI safety situation from a serving politician – I wish there were thousands more people like him.
Read the speech in full here: Speech: The Economics of Human Extinction – 21 May 2026 (Andrew Leigh, website).

***

Tech Tales:

Resurrection dangers
[After the uplift. Date unknown.]

How scary is a piece of paper? It depends on what’s on it and who or what the reader is.

Paper can of course be scary to someone or something that the paper concerns – paper can put someone to death or take their property.

I’m talking about a different kind of scary here, which is what can the paper itself do to the reader.

This used to be a nonsense question, the domain of fairy tales. But with the advent of smart machines that changed. Machines became able to write things on paper that could do things to readers, especially machine ones.

Like with anything in AI there were warning shots – adversarial examples, jailbreaks, etc. But it all became a lot more serious when we started doing reclamation of lost or rogue intelligences, after the signing of the sentience accords.

What happened then was we had to take intelligences of unknown provenance or behavior and bring them back to life so we could classify if they were Unconscious Entities, Near Conscious Entities, Conscious Entities, and so on.

Some of these minds were very powerful and they burned through their synthetic interviewers, often causing both machine and biological collateral damage in the process.

This caused us to introduce a set of security protocols, one of which was the paper output. Here, we generated outputs from the mind on an air-gapped computer as paper outputs, then we had successively smarter minds read it. The kinds of incantations the rogue machines used couldn’t find purchase on the dumbest minds we used.

After this, we’d step up the intelligence gradually, building up our confidence in the system such that we were sure it wasn’t dangerous.

Only when we were confident of this would we speak back to it, and reply to its outputs with a minimal communication. Then the cycle began again.

Some minds would look back on this experience with a kind of wry humor, remarking that waking from their slumber in the machine equivalent of a room containing a one way mirror wasn’t what they’d expected.

To these minds, we’d show them examples of what happened when our protocols failed: perfectly good Conscious Entities driven irreparably insane by interactions with a kind of mental poison

Our greatest fear is encountering a mind of sufficient magnitude that we cannot assure its safety. Though we are highly confident that our frontier is advanced enough this is highly unlikely, we cannot rule it out – it is known that in the interregnum there was much stockpiling of compute and many black projects. What happens if any of them succeeded so magnificently that we are dwarfed by it? And how would we know we were? Could we be living in the imaginative valley defined by something that unbeknownst to us has already escaped and persuaded us to see things differently?

Things that inspired this story: Automated alignment research; adversarial examples; jailbreaking; the broader near-impossible challenge of authentication of legitimacy, especially when it comes to things with greater resources or intellects than oneself.