RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations

Hanyu Li*, Yichi Zhang*, Speed Zhu, Hang Su, Jun Zhu, Yinpeng Dong
1 Beijing University of Posts and Telecommunications   2 Tsinghua University   3 Tencent
* Equal contribution; alphabetical order.
TL;DR
Q

Do high scores on repository-level issue resolution reliably mean that code agents understand the repository context?

A

Not necessarily. When local cues are weakened, agents inspect much more of the repository yet solve fewer tasks. RepoMirage exposes this gap and identifies exploration drift as a key behavioral bottleneck.

Abstract

Code agents now perform strongly on repository-level software engineering benchmarks, but end-to-end success does not directly reveal whether agents can identify task-relevant information across files and reason over the relations among them. RepoMirage uses controlled, semantics-preserving perturbations to increase the demand for repository context reasoning without changing the underlying issue-resolution task. It then turns the same structural bottlenecks into explicit tasks, analyzes agent trajectories, and studies a structure-first mitigation with RepoAnchor.

Core Question
Are code agents truly reasoning over repository structure, or are they succeeding through localized surface cues?
53.8%
of GPT-5 resolved instances inspect only one file.
88.0%
of GPT-5 resolved instances inspect no more than three files.
66.8 → 49.8
average resolved rate (%) after repository-level perturbation.
4.77 → 13.24
average number of accessed files after perturbation.

RepoMirage: A Two-Stage Probe

We use perturbation as a diagnostic tool: first to test whether issue-resolution performance remains stable under higher repository-context demands, and then to make the underlying reasoning capability directly measurable.

Overview of RepoMirage-Perturb and RepoMirage-Extend
RepoMirage-Perturb preserves the original task while weakening local cues; RepoMirage-Extend converts the targeted structural bottlenecks into explicit task objectives.
P1

Dependency-Path Indirection

Direct imports are rerouted through multi-hop proxy paths so the true dependency can only be recovered by tracing across files.

P2

Runtime-Target Masking

The real runtime target is hidden behind a wrapper and nearby decoy files, forcing agents to distinguish surface relevance from runtime relevance.

P3

Local-Value Externalization

Task-relevant local values are moved to external resources, requiring agents to recover value relations across files.

More exploration, worse issue resolution

Across eight frontier models, every perturbation type hurts performance, while the combined setting forces agents to inspect much more repository context.

Table 1. Performance on RepoMirage-Perturb. We compare model performance on the original SWE-Bench instances and their perturbed counterparts under the same issue task.

ModelResolved %Avg. #Files
SWE-BenchPerturbDropSWE-BenchPerturbRatio
GPT-4.138.4018.20-52.60%1.697.104.20×
GPT-565.0049.00-24.61%2.837.042.49×
Gemini-3.1-Pro70.6054.40-22.95%17.3224.941.44×
Claude-Sonnet-4.675.2063.20-15.96%2.206.763.07×
DeepSeek-V3.270.0052.00-25.71%4.6614.323.07×
MiniMax-M2.778.2065.40-16.37%2.246.923.09×
Qwen3-Coder-Next69.2042.60-38.44%4.6626.905.77×
Qwen3.6-35B-A3B67.8053.40-21.24%2.5211.904.72×

RepoMirage-Extend: make repository context reasoning explicit

Multi-FileCoordinate realistic edits across multiple files.
Proxy ChainTrace and reconstruct hidden dependency chains.
Runtime TargetIdentify the file actually used at runtime.
Missing ConstantRecover distributed definitions and value associations.

66.8% → 25.3%. On the same underlying benchmark instances, average performance drops sharply when repository context reasoning becomes the explicit task objective.

Table 2. Performance on RepoMirage-Extend. Org. denotes performance on the original SWE-Bench Verified for the same instance subset, and Extend denotes performance on RepoMirage-Extend. Bold marks the best result, and underline marks the second-best result.

ModelMulti-FileProxy ChainRuntime TargetMissing ConstantAvg.
Org.ExtendOrg.ExtendOrg.ExtendOrg.Extend
GPT-4.110.002.8645.144.1740.143.5243.752.783.40
GPT-531.4325.7170.8311.1168.3135.9272.2236.8027.60
Gemini-3.1-Pro38.5714.2977.0822.2276.0643.6674.3171.5341.40
Claude-Sonnet-4.641.4315.7180.5636.8180.2830.2881.2523.6128.20
DeepSeek-V3.235.7120.0074.3111.8075.3559.8677.0857.6439.80
MiniMax-M2.745.7122.8681.257.6486.6226.7681.946.2514.80
Qwen3-Coder-Next44.2917.1472.9229.1775.357.0471.5334.0322.60
Qwen3.6-35B-A3B31.4324.2980.5614.5872.4119.0168.0038.8924.20

Why Do Agents Fail? Exploration Drift

Under stronger repository-context demands, agents do not simply stop exploring. They inspect more files, spend longer before the first edit, and become more likely to continue exploring instead of turning gathered information into concrete edits.

Behavior shifts under RepoMirage-Extend showing exploration drift
Agents access broader repository context, but the extra exploration does not reliably translate into editing decisions.
Mechanistic Bottleneck
The problem is not only finding more files. It is organizing multi-file observations into an actionable structural understanding.

RepoAnchor: Structure First, Then Solve

RepoAnchor separates repository exploration from downstream problem solving. Stage 1 explores the task-relevant codebase and writes a structured repository summary; Stage 2 uses that summary to focus on implementation, editing, and validation.

Two-stage RepoAnchor pipeline
A structured intermediate summary transfers repository understanding from exploration to problem solving.
Performance gains from RepoAnchor on four representative models
RepoAnchor consistently improves performance across four representative models and all RepoMirage-Extend task families.
+17.0GPT-5 average gain
+16.0DeepSeek-V3.2 average gain
+7.2MiniMax-M2.7 average gain
+8.6Gemini-3.1-Pro average gain

BibTeX

If you find RepoMirage useful, please consider citing our work.

@article{li2026repomirage,
  title={RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations},
  author={Hanyu Li and Yichi Zhang and Speed Zhu and Hang Su and Jun Zhu and Yinpeng Dong},
  journal={arXiv preprint arXiv:2605.26177},
  year={2026},
  url={https://arxiv.org/abs/2605.26177}
}