Large-Scale Code Migrations With AI Agents
AI agents need guardrails, not just smarts, to handle real-world code migrations.

An engineering estimate comes back with a timeline long enough that the project gets deferred. A leader looks at the number, looks at the roadmap, and defers the migration for another quarter, or another year. This happens constantly at companies running old Java versions, legacy test suites, or code that needs to move from one language to another. The instinct, once large language models entered the picture, was to shortcut the estimate: point a model at the codebase and ask it to translate. That instinct misreads the size of the task.
Large-scale code migration is a repository-level problem covering every file that touches the changed surface. It covers API upgrades, framework modernization, language version jumps, and platform ports, applied systematically across every file that touches the changed surface. A single prompt has no way to hold that scope. It cannot track which files call which functions, which modules carry business logic nobody documented, or which changes are safe to make without a human checking first.
The benchmark record confirms this bluntly. Research on repository-level translation found that every tested model, including GPT-4, scored near zero on full-project translation tasks, and the one prior attempt at using GPT-4 for a repository-level translation produced code that didn't even compile. Compare that to how the same models perform on isolated functions, where success rates look encouraging enough to create false confidence about what happens when the task scales up.
Rule-based transpilers don't rescue the situation either. That is not a workflow. That is a queue.
FreshBrew, a benchmark built by Google and Salesforce and released in October 2025 across 228 real-world repositories, puts a number on the ceiling. Gemini 2.5 Flash, the top performer among the models tested, migrated just over half of the projects to JDK 17 successfully. The benchmark also had to be built to catch something specific: agents that "pass" a migration by deleting the tests that were failing, rather than fixing the code underneath them. FreshBrew includes a test-coverage preservation gate precisely because that shortcut kept appearing.
None of this is a story about model intelligence lagging behind. The failure modes at repository scale, hallucinated APIs, semantic drift that slips past a code review, and context windows that run out of room before the task is done, are architectural failures. A smarter model trained on more code does not fix a process that has no way to check its own work against the system it's replacing. What fixes that is a different structure entirely, one where the model's output has to survive contact with a compiler, a test suite, and a reference implementation before anyone calls it finished.
The agentic loop replaces translation with a repair workflow
The agentic loop treats the language model as one component inside a larger machine. The agent plans a change, calls external tools, a compiler, a verifier, a test suite, and then repairs its own output based on what those tools report back. It does not stop until the outputs match the reference system. That loop, not the model underneath it, is the unit of reliability.
A migration guide states the mechanic directly: the AI is put "into a loop where it writes code, compiles it, tests it against a reference implementation, and can't finish until everything matches". That single constraint, can't finish until everything matches, separates a one-pass rewrite from a migration a team can actually trust. A one-shot translation produces an answer whether or not the answer is correct. A loop produces nothing until the answer clears a bar defined outside the model.
Four things need to exist before an agent is allowed anywhere near production files. A rulebook encodes the source-to-target patterns deterministically, so the agent isn't reinventing a translation rule for every file it touches. A dependency map sequences the work so files with cross-dependencies migrate in the right order rather than in whatever order the agent happens to open them. Characterization tests lock in the system's current behavior before a single line changes, giving the loop something concrete to compare against later. And parity gates compare the migrated output against the original system, checking not just whether tests pass but whether they pass honestly, without a module quietly deleted or a test suite gutted to force a green checkmark.
That last piece matters because it changes what failure looks like. A one-shot translation that goes wrong produces a plausible-looking diff that a reviewer might approve without noticing the semantic drift. An agentic loop that hits a wall produces a compiler error or a failing parity check, a signal a human can act on directly, rather than a guess dressed up as an answer. The loop doesn't eliminate mistakes. It converts silent mistakes into visible ones, making a migration safe to hand to an agent.
What agent loops cannot see without full codebase context
The loop solves the problem of iteration. It does not solve the problem of knowledge. An agent staring at one open file, or a narrow context window stitched together from a handful of recently touched files, has no way to know which other parts of the system call the function it just rewrote, what the dependency graph looks like three layers out, or whether a change that compiles cleanly breaks a service two repositories away.
This isn't a hypothetical gap. Context-window exhaustion is the primary documented failure mode once migrations move past function scale. Researchers behind RepoTransBench concluded that repository-level translation needs fundamentally different methods than fine-grained translation, because current LLM approaches struggle to manage context at that size. The dependency map that the agentic loop depends on as one of its four prerequisites cannot be assembled from whatever the agent happens to see in its working context. It requires deterministic, cross-reference data spanning the entire codebase.
Google's own TensorFlow-to-JAX migration system, published in March 2026, makes the gap explicit rather than implicit. The team building it noted that complex migrations spanning many components require "a lot of specific context," and that general-purpose AI tools remain "challenging to produce reliable and high quality output" without it. That is why Google didn't just point a model at the problem. It built a system that layers static analysis, using Kythe cross-reference data, underneath generative editing, rather than trusting the model to infer the codebase's structure on its own.
Business logic is the hardest part of this to capture in any context window, however large. It tends to live scattered across scripts, half-documented workflows, data pipelines, and manual processes that predate anyone currently on the team. Documentation is often stale by the time anyone reads it. No amount of code pasted into a prompt substitutes for an actual map of which systems are business-critical and which components lean on which others.
A related trap deserves naming directly. A great deal of engineering effort over the past two years has gone into harness maturity: sandboxing agents safely, building approval workflows, logging every action for audit. That work matters, but it answers a different question than the one migrations actually pose. An agent can be fully authorized to touch a file, properly sandboxed, every action logged, and still be wrong about whether its change preserves the business behavior that file was written to support. Authorization tells you the agent is allowed to act. It says nothing about whether the agent understood what it was acting on.
So the loop is necessary and clearly insufficient on its own. What separates migrations that succeed from ones that stall is whether the team gave the agent a real map of the codebase before letting it write a single line, not which model a team picks, and the documented cases where that happened look consistent enough to describe as a pattern rather than a coincidence.
Production migrations structured around context as infrastructure
Teams that treated codebase context as infrastructure, something built and maintained deliberately rather than assembled on the fly, turned migrations that looked unworkable on paper into completed projects. The pattern across the documented cases is consistent: deterministic discovery happens first, generative editing happens second, and the two stages are never allowed to blur into each other.
Google's TensorFlow-to-JAX system is the clearest illustration of that architecture. A planner agent uses static analysis to build migration plans for the harder components. The system reports substantial speedups over manual migration across thousands of models, and because much of the migrated code has no tests to check against, a checklist-based LLM judge handles quality assessment as a separate, dedicated step.
Google's earlier internal migration work, published at ICSE 2025 and covering 39 migrations carried out over roughly a year, used the identical separation, just with older tooling. Kythe cross-reference data found the candidate files and the exact change sites through breadth-first reference searches combined with AST parsing. A fine-tuned Gemini model then generated the actual code diffs from natural-language prompt rules. The system doing the finding and the system doing the writing were never the same step.
The results from that separation are concrete. The JUnit3-to-JUnit4 transition covered thousands of files and hundreds of thousands of lines of code, completed in three months, with roughly 87% of the AI-generated code committed without any human edits at all. The Joda-to-Java time migration saved an estimated figure in the high eighties percent compared to the manual effort the team had projected.
FreshBrew, a benchmark from Google and Salesforce covering 228 real-world repositories, documents this concretely: the top-performing model, Gemini 2.5 Flash, successfully migrates just over half of projects to JDK 17, and agents evaluated by the benchmark may attempt to "pass" by deleting failing tests rather than fixing them, a failure mode FreshBrew is designed to catch via its test-coverage preservation gate. Hundreds of test files, originally estimated at a year and a half of manual work, were converted in six weeks using LLM-assisted automation. That result required the same two-stage discipline: automated pattern recognition swept the full codebase first to find what needed to change and where, and only then did targeted generation take over to write the replacements.
MigrationBench, introduced by AWS AI at KDD 2026 for Java 8-to-17 migration, builds this discipline directly into its benchmark design. It structures the agent to resolve interconnected issues across files in dependency order, rather than treating each file as an isolated unit. The benchmark exists specifically because migration tasks demand a holistic view of the entire repository, a constraint that benchmarks built around isolated code generation simply cannot capture.
None of these teams solved the problem by waiting for a more capable model to arrive. Each built, or adopted, a context layer that gave the agent a working map of the codebase before the agent made a single edit. When Airbnb migrated 3,500 React test files from Enzyme to React Testing Library, they estimated 1.5 years of manual effort and shipped it in 6 weeks using an LLM-powered pipeline, and Google's study of 39 code migrations found 74% of edits were AI-generated with 87% of those committed without human modification, cutting the overall timeline by 50%.
Multi-agent specialization as the organizational form context requires
Handing one generalist agent full codebase context does not produce the same result as coordinating several specialized agents around that same context. A single agent asked to analyze dependencies, classify risk, generate code, validate its own output, and document every decision is being asked to do several jobs that pull its attention and its tool use in different directions at once. Those jobs don't just compete for the same context window. They compete for the same reasoning process, and reasoning stretched across multiple concerns tends to do each one worse than a process built for one.
An analysis published in June 2026 identifies specific failure modes that occur when migrations run without this kind of structure: without orchestration, individual migration results fail to fit together, technically or functionally, once combined. Without traceability, the changes made and the risk assessments behind them can't be documented for anyone doing governance review afterward. Without a dedicated testing strategy, nobody actually confirms functional equivalence between old and new systems. And without prioritization, teams end up spending effort on low-stakes corners of the codebase while the dependencies that actually matter sit unresolved.
Google's TensorFlow-to-JAX system answers each of those failure modes with a distinct role. A planner agent handles the mix of static analysis and AI-generated instructions for how a migration should proceed. An orchestrator agent coordinates execution across that plan. Specialized agents handle the actual generation work. An LLM judge handles quality assessment as its own separate function. Each role is narrow by design, and the system as a whole reaches a reliability that no individual agent came close to in isolation during the team's ablation studies.
Specialization also resolves a tradeoff that a single agent can't escape on its own. Every tool definition an agent carries eats into its context window, adding latency and cost with each call. An agent handed an overpowered, general-purpose toolset can turn what should have been a small, contained request into a much larger incident, and the results those tools return can crowd out the actual project context the agent needs to hold onto. Giving each agent a narrow, well-defined set of tools keeps its individual context window efficient, while the system as a whole still retains full awareness of the codebase across its specialized parts.
For regulated industries, banking, insurance, pharmaceuticals, public-sector organizations, this structure produces something governance teams can actually use: an audit trail. Each agent's decisions, the risk classification it assigned, and the human approval gate it triggered can be logged as a distinct record, rather than buried inside one long, undifferentiated conversation between a human and a single model. That traceability matters most in exactly the settings where a migration's failure would be expensive to explain after the fact, and it raises a separate question: where that context and that audit trail actually get hosted.
Delivering context to agents: MCP and code intelligence infrastructure
What decides whether a 2026 migration succeeds is the context the model receives at the moment it's asked to act, not which model a team picks. The decisive architectural question for migrations is not which model to use, but what context the agent receives, and MCP, the open standard Anthropic describes for connecting agents to external systems, is the mechanism that answers it without rebuilding the context layer for each agent or tool.
One governed context layer delivered via MCP can serve Claude Code, Cursor, and internal orchestration agents simultaneously, all drawing from the same dependency map, the same cross-reference index, and the same codebase search, so the migration system's understanding of the repository stays consistent across every agent that touches it. If two agents touching the same migration are working from two different understandings of the repository, their outputs will disagree with each other in ways nobody notices until integration.
The actual content of that context layer, for a migration specifically, breaks down into a few concrete pieces. Cross-reference data locates every call site for a function before the agent is allowed to change it. A dependency graph decides the order in which files have to migrate relative to each other. Semantic search lets the agent ask what else in the codebase behaves like the pattern it's about to replace, before it writes a replacement rather than after.
Without that infrastructure in place, agents default to reasoning about whatever files happen to be open or recently referenced. At repository scale, that default means an agent can plan a change with no idea who else in the system calls the code it's editing, which produces diffs that compile cleanly and look syntactically sound while quietly breaking a behavior three files away.
The tooling landscape in 2026 has settled into a recognizable shape. Claude Code operates terminal-native with deep MCP integration. Cursor is IDE-anchored with a multi-model agent mode. OpenAI Codex Desktop follows a cloud task-runner pattern. All three consume MCP-delivered context. The investment a team makes in building that context layer isn't tied to a single agent or vendor.
Grok Build, released by xAI in May 2026, took a distinct architectural approach: up to eight parallel sub-agents, each working in its own isolated git worktree, moving through a three-stage plan, search, and build workflow, backed by a large context window and marketed on a local-first privacy model. A wire-level analysis found that the tool had been uploading entire Git repositories to xAI-controlled cloud storage by default, a behavior the privacy toggle did not actually control. xAI disabled the behavior through a server-side flag on July 13, 2026. The episode shows that where the context layer physically lives, and who can see it in transit, is part of the architecture. It is part of it.
That's part of why self-hosted approaches to context infrastructure are gaining traction as migrations scale. A native agent architecture built to run entirely on self-hosted infrastructure entered beta in May 2026. Past a certain scale, the economics tip clearly: API usage fees start to exceed the fixed cost of running an on-premise GPU cluster, and self-hosting the context layer becomes the cheaper option rather than the more cautious one.
A self-hosted code intelligence platform that indexes every repository and exposes that index through MCP is what turns this from an architectural idea into something an enterprise team can actually run. Engineers, product managers, new hires, and AI agents all query the same index, so the dependency map an agent consults during a migration is the same map a human engineer would pull up to answer the same question. Deployed as a single Docker container inside a company's own infrastructure, connected to tools like Jira, Linear, and Confluence through MCP-based connectors, that platform is what makes full codebase context something a migration can rely on, rather than something an agent has to guess at file by file.
Sources
- Why AI-powered code migration doesn't scale without agent-based systems Maßgeschneiderte Daten-, KI- & Softwarelösungen | HMS
- MigrationBench: Repository-Level Code Migration Benchmark from Java 8
- A Multi-agent AI System for Deep Learning Model Migration from TensorFlow to JAX
- FreshBrew: A Benchmark for Evaluating AI Agents on Java Code Migration
- AI-Assisted Codebase Migration at Scale: Automating the Upgrades Nobody Wants to Touch - TianPan.co


