First things first: what is an AI agent?

A typical chatbot works on a "you ask, it answers" basis. An AI agent, by contrast, takes action on its own: it can install software, call tools and run code, then pull the results together and report back to you. What Paper2Agent does is turn a single paper into this kind of agent.

From passive documents to agents that get things done

Most papers are meant to be read, not used. To apply a method described in a paper, researchers usually have to install the authors' software first and work out how to feed it data, and that step alone can leave them stuck for days.

As Earth.com reported on October 4, Paper2Agent, developed by Stanford postdoctoral scholar Jiacheng Miao, associate professor James Zou and colleagues, skips that step: users ask questions in plain language, and the agent answers by running the paper's own code. The research has been published in the journal Nature.

The process has three stages, each handled by its own agent:

  • Set up the environment: the first agent installs the software environment the code needs.

  • Build the tools: the second agent packages the paper's main methods as tools that other AI systems can call.

  • Check the results: the third agent tests each tool against the results and figures reported in the paper, and any tool that keeps failing is dropped.

Paper2Agent's three-agent workflow: set up the environment, package the tools, test them against the paper's results, then bundle everything into an MCP package
Paper2Agent's three stages: set up the environment, wrap the methods as tools and test them against the paper's results; the tools that pass are bundled into an MCP package (AI-generated illustration)

The tools that pass are packaged with the Model Context Protocol (MCP), together with the paper's text, its data and common analysis steps. MCP is a shared standard that lets AI access external tools and data; think of it as a universal power outlet for the AI world. Zou says an MCP lets AI represent a paper in a form that is easy for agents to access, "almost like a filing system."

What the numbers say

Take AlphaGenome, Google DeepMind's genome model. The system built 22 tools for it in about 45 minutes with no human stepping in at all. It ran on a personal laptop, and the whole job cost just US$14.

On 30 open-ended research questions that required chaining several tools, the agent scored 82.7%. Claude working directly from the code managed only 56.7%, while Biomni, another AI research assistant, reached 72.2%.

In a larger-scale test, 74 of 100 computational biology papers were successfully turned into agents, and 593 of the 599 tools the system proposed passed testing. On 300 questions, the agents scored 91.2%, ahead of the 80.3% achieved by working directly from the code. The failures came mainly from missing code, missing data or software that wouldn't install.

Paper2Agent test results: 82.7% success rate, 74 of 100 papers converted, 593 of 599 tools passed, 22 tools built in about 45 minutes, US$14 cost
Key numbers: 82.7% on open-ended questions, 74 of 100 papers turned into agents, and 22 tools built for AlphaGenome in about 45 minutes for US$14 (AI-generated illustration)

Three agents team up to find a psoriasis gene

The most striking experiment had three paper agents divide up the work. First, the AlphaGenome agent predicted that in CD4+ T cells (a type of white blood cell that helps regulate immune responses), a psoriasis-linked spot in the genome affects a gene called GPR137 more than any other nearby gene.

The researchers then connected agents built from two earlier papers to cross-check the data. Of five candidate genes, only GPR137 matched, and only in cells that had been stimulated. The whole process involved no new experiments; every check used existing data.

The limits are clear too

The authors stress that the agents are not independent scientists: researchers still choose the direction, and the evidence still has to be weighed by people. Anyone who publishes an agent has to keep maintaining the software underneath it, and questions of security, ownership and academic credit remain unresolved. The authors also suggest that journals may one day ask papers to include an "agent availability" section.

What it means for Taiwan

For research institutions, the most telling part of Paper2Agent is not its accuracy but the 26 papers that failed: missing code, missing data, environments that wouldn't install.

Taiwan's research institutes and universities produce large volumes of project reports and experimental data every year, but code and data are often scattered across personal computers and left unmaintained once a project closes. If knowledge management stops at "archiving PDFs," it won't be able to plug into the agent-based ways of working that lie ahead.

On the left, an old file room piled high with PDFs and cardboard boxes; on the right, a lab where code, data and environments are neatly in place and AI agents can work with them directly
Knowledge has to run, not just be filed: only when code, data and runtime environments are delivered together can research results plug into agents (AI-generated illustration)

A practical first step comes down to two things:

  • Deliver code, data and runtime environments as part of the research output, not just a report.

  • Set the rules on access to internal data, maintenance responsibilities and credit attribution first, before moving on to making reports "speak."

Sources