Engineering Context

Automated Dependency Upgrades Across Multiple Repositories

Coordinating dependency updates across repositories prevents breaking changes and review overload.

Staff Writer · · 10 min read
Cover illustration for “Automated Dependency Upgrades Across Multiple Repositories”
AI Coding Agents · October 5, 2026 · 10 min read · 2,275 words

A dependency upgrade in one repository is a local decision. Across dozens or hundreds of repositories, that same upgrade becomes a coordination problem, and the risk that matters most is no longer the gap between dependency versions but the gap between repositories themselves. A single repository can track its own dependency tree, run its own tests, and decide for itself when a new version is safe to adopt. Nothing about that arrangement scales, because each repository drifts on its own schedule, with no shared clock and no shared view of what every other repository is doing with the same library.

That drift compounds in a specific way. The research literature on breaking changes in software ecosystems treats this as a structural property of dependency networks: breaking changes propagate through them, and downstream consumers fail at a distance from the change that caused the failure.

J.P. Morgan AI Research's work on LLM agents for dependency upgrades states the industrial version of this problem directly: thousands of code repositories in industry use a given framework, and when that framework is upgraded, only some of the repositories that need the required code updates actually receive them. A per-repository view cannot see this, because the dependency graph that matters spans repository boundaries, and no single repository's CI pipeline has visibility into any other repository's consumption of the same code.

Two distinct failure modes follow from this. The second is breaking-change propagation: a dependency bump can pass every test in the repository where it was made and still break a downstream consumer in a repository that was never touched, never tested against the new version, and had no way to know the change was coming. Both failure modes share a root cause. They are failures of coordination across repositories, and no amount of diligence applied one repository at a time closes that gap.

What PR overload and interrupt-driven review cost engineering teams

The intuitive fix, running a dependency bot on every repository independently, multiplies the demands on reviewers faster than it reduces their manual work. If there's no layer that coordinates across repositories, automation multiplies the number of things demanding review faster than it reduces the manual work those reviews replace.

The arithmetic behind this is straightforward. Those deprioritized PRs age and go stale, and the dependency debt they were supposed to retire doesn't disappear. It reappears as a backlog of open, unreviewed pull requests sitting in every repository at once. GitHub's own account of this problem is explicit: before its February 2026 Dependabot update, upgrading a single dependency across three directories in one repository produced three separate pull requests, and the cross-directory grouping feature exists because that kind of sprawl was the primary operational complaint from teams running Dependabot at scale.

The deeper cost sits in how these PRs interrupt work, not just how many of them exist. The attention of the people whose attention is most expensive and most load-bearing elsewhere in the organization gets taxed by this. A process meant to save engineering time instead consumes the time of the engineers best equipped to spend it on harder problems.

The architectural split between Dependabot and Renovate

Two tools dominate automated dependency management, and they are built on different premises, not competing to do the same job better. Dependabot is built directly into GitHub and optimizes for activation with almost no setup. Renovate (distributed commercially as Mend Renovate) optimizes for expressing a dependency policy once and having that policy apply consistently across a large and heterogeneous repository estate. Choosing between them is a question of which premise matches the shape of the organization doing the choosing.

Dependabot's advantage is that it requires minimal configuration to start covering GitHub-hosted repositories, which makes it the natural choice for teams already standardized on GitHub who want coverage running the same day with little setup overhead. Its February 2026 update, the cross-directory grouping feature available on github.com and shipping in GHES 3.21, addresses the monorepo PR sprawl described above directly: configuring directory groups in a repository's dependabot.yml file consolidates updates for the same dependency into one pull request, regardless of how many directories that dependency touches.

Renovate's advantage becomes visible in organizations that are not fully standardized on one platform or one policy, and its CLI covers more than 90 package managers, including npm, Java, Python,.NET, Scala, Ruby, Go, Docker, and Terraform, putting its ecosystem coverage on par with Dependabot's.

The differentiator that tends to decide the choice for enterprise teams is Renovate's shared-preset model. Configuration resolves in five layers, built-in defaults, global config, inherited config, resolved presets, and finally repository config, with packageRules applied last within each layer. Mend Renovate Enterprise builds on top of the free Community edition with Merge Confidence workflows, reporting APIs, self-hosted deployment, and paid support with service guarantees; Merge Confidence ratings themselves are available across tiers, while the workflows that act automatically on those confidence levels are an Enterprise-tier capability.

Neither tool is the correct answer in the abstract. The decision follows from the organization's actual platform footprint and the actual complexity of its policy.

Configuring grouping, scheduling, and auto-merge to manage PR volume

Once a tool is chosen, the configuration choices that matter most have nothing to do with which packages get updated. They concern how updates are batched into PRs, when those PRs are generated, and which of them get merged without a human in the loop. These three levers, grouping, scheduling, and auto-merge, determine together whether dependency automation quiets the review queue or floods it.

Grouping is the first lever. Related updates, all AWS SDK packages, or all React ecosystem packages, for instance, should land in a single pull request so a reviewer evaluates one coherent change instead of working through a parade of individual version bumps that all touch the same underlying system. Mend Renovate Enterprise extends this with Merge Confidence workflows that group by confidence band: high-confidence updates travel together on a fast path toward merge, while lower-confidence updates get separated out for the closer review they actually need. For monorepos specifically, Dependabot's cross-directory grouping collapses what used to be one PR per directory per dependency into a single PR per dependency across every directory it touches, configured through directory groups in dependabot.yml.

Scheduling is the second lever. Dependency scans should run on a cadence that respects the team's own release calendar, not on a cadence that happens to trigger PRs during a release freeze or a high-deployment window. Renovate's scheduling module supports configurable intervals and can respond to repository events through webhooks, which allows priority updates, a security patch, for example, to be processed outside the normal schedule when that is warranted.

Auto-merge is the third lever and the one that converts the first two into an actual reduction in reviewer load. Auto-merge is safe to enable when it is gated on two conditions at once: CI has to pass, and the update has to clear a confidence threshold, whether that threshold is defined simply as patch-only changes or expressed as a Merge Confidence score above a configured level. The effect is to remove the reviewer interrupt entirely for the large, routine category of patch-level updates, and preserve human attention for the smaller category of changes that actually carries risk.

Cross-repository visibility as a prerequisite for catching breaking changes before PRs are opened

No amount of grouping, scheduling, or auto-merge configuration solves the problem described at the start of this article, because none of those levers give a tool visibility into repositories it isn't running on. A breaking change made in one repository and detected by that repository's own CI tells a team nothing about which other repositories consume the changed code and how.

The mechanism is specific. A field or interface gets updated inside a library repository. That library is consumed by several downstream service repositories, maintained by different teams, none of which were party to the update. The tool managing the upgrade sees the library's own change in isolation, runs the library's own tests, and opens a PR that passes, because passing is all it can tell from its own repository. The break only becomes visible when a consuming service tries to integrate against the new version, at which point the failure is disconnected in time and in ownership from the change that caused it.

J.P. Morgan AI Research's multi-agent framework was built to close exactly this gap. The Summary Agent's specific job is to map where a breaking change lands across the codebase before any code gets modified, which reframes the upgrade task: the first question is where the break will occur, not how to fix it.

That reframing has a direct operational consequence. Before opening upgrade PRs across a repository estate, a team needs an answer to a question no single repository can answer on its own: which other repositories consume this dependency, and how do they consume it? Answering that question requires indexing every repository in the estate, not scanning the one repository where the upgrade originates. That dependency graph isn't fixed, either. New consumers of a given library appear continuously as other teams add dependencies of their own, so the map has to be maintained on an ongoing basis to stay accurate rather than built once and trusted indefinitely. Code search and navigation tools capable of indexing an entire repository estate, answering a query like which services import this interface, make this tractable without requiring any single engineer to hold the whole dependency graph in their head.

LLM agents and the upgrade workflow for breaking changes

Patch and minor version bumps are well served by a properly configured PR bot. Major upgrades that require actual code changes to restore compatibility across multiple repositories sit outside what a PR bot can do, because a PR bot can open a change but cannot write the fix a breaking change demands. That gap is where LLM-based agents become the practical option rather than a speculative one.

J.P. Morgan AI Research tested its framework on an industrial Java upgrade scenario built from three synthetic repositories containing major breaking changes, and the framework reached a precision of 71.4% while using fewer tokens than prior approaches to the same task. That figure demonstrates both that agent-driven upgrades are workable and that they are not yet complete: a precision in the low seventies means a meaningful share of the recommended fixes still need human correction, and teams adopting this approach should plan for that review step rather than treat agent output as final.

The three-agent architecture divides the work along clean lines. The Summary Agent maps where a breaking change propagates across the codebase. The Control Agent sequences the remediation work and manages the dependencies between individual fixes, so that fixes get applied in an order that doesn't break on itself. The Code Agent implements the recommended changes file by file. Each agent's output bounds the next agent's task, which is part of why the architecture holds together across repositories rather than collapsing into a single undifferentiated code-generation step.

An agent's usefulness is capped by what it can see. An agent limited to the file it is currently editing has no way to reason about whether that edit breaks a consumer living in a different repository, so full codebase context is a precondition for an agent to operate safely on this kind of task, not a convenience layered on top of it. Agent-native workflows built around tools like Claude Code and Cursor are well suited to work that is bounded, testable, and reviewable, and dependency upgrades fit that description closely: the success criterion is unambiguous (CI passes, nothing regresses), the scope of the task is defined by the upgrade itself, and the output arrives as a pull request that a human can still read and judge. The Model Context Protocol (MCP) is becoming the standard way to hand agents the external tools, data, and context they need at runtime, and exposing cross-repo search results, symbol definitions, and usage graphs to agents through MCP connectors is what supplies the organizational knowledge an agent cannot derive from a single file on its own.

Diagram: Three Agents, One Upgrade: How the Multi-Agent Framework Divides the Work. Visualizes: Visualize the three-agent pipeline from J.P.

Security and data residency requirements that shape how automation is deployed in enterprise environments

None of the preceding architecture, tool choice, configuration, cross-repo visibility, or agent-driven remediation, can be acted on inside an enterprise until its security and data residency implications are settled. Code isolation, data residency, and the scope of an agent's privileges are not special requirements reserved for regulated industries. Every enterprise security team applies them as the baseline before it lets any agent touch production code.

Two distinct risks apply here, and they call for different controls. The first arises at inference time: sending code to a cloud LLM provider means that code leaves the organization's network, full stop on the simple version of the concern. Enterprises with on-premises or air-gapped requirements need to evaluate this provider by provider rather than assume uniform treatment.

The second risk arises at runtime, once an agent has write access to repositories and to CI systems. An agent holding that kind of access is a privileged actor, and it can combine individually legitimate permissions into actions nobody specifically authorized. The appropriate response is to scope an agent's access to the minimum required for the specific upgrade task in front of it, and to require human review before merge on any change that touches an interface shared across services. That second condition matters most precisely where this article began: at the boundary between repositories, where a change that looks safe in isolation is exactly the kind of change most likely to break something the agent, and the automation surrounding it, cannot see.

Sources

  1. Dependabot can group updates by dependency name across multiple directories - GitHub Changelog
  2. GitHub - renovatebot/renovate: Home of the Renovate CLI: Cross-platform Dependency Automation by Mend.io · GitHub
  3. LLM Agents for Automated Dependency Upgrades
  4. renovate/docs/usage/key-concepts/automerge.md at main · renovatebot/renovate
  5. Breaking Changes in Software Ecosystems: A Systematic Literature Review
Filed underAI Coding Agents

More in AI Coding Agents