Jimenez et al. [2024]
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan.
SWE-bench: Can language models resolve real-world GitHub issues?
In Proc. ICLR, 2024.
Yang et al. [2024]
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press.
SWE-agent: Agent-computer interfaces enable automated software engineering.
In Proc. NeurIPS, 2024.
Wang et al. [2024a]
X. Wang, Y. Chen, L. Yuan, Y. Zhang, Y. Li, H. Peng, and H. Ji.
Executable code actions elicit better LLM agents.
In Proc. ICML, 2024.
Yao et al. [2023]
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao.
ReAct: Synergizing reasoning and acting in language models.
In Proc. ICLR, 2023.
Wang et al. [2024b]
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, et al.
A survey on large language model based autonomous agents.
Frontiers of Computer Science, 18(6), 2024.
Hong et al. [2024]
S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Lin, et al.
MetaGPT: Meta programming for a multi-agent collaborative framework.
In Proc. ICLR, 2024.
Wu et al. [2023]
Q. Wu, G. Banber, Y. Zhang, Y. Wu, B. Li, Z. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang.
AutoGen: Enabling next-gen LLM applications via multi-agent conversation.
arXiv preprint arXiv:2308.08155, 2023.
Fourney et al. [2024]
A. Fourney, G. Bansal, H. Mozannar, C. Tan, E. Horvitz, et al.
Magentic-One: A generalist multi-agent system for solving complex tasks.
arXiv preprint arXiv:2411.04468, 2024.
Schick et al. [2023]
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom.
Toolformer: Language models can teach themselves to use tools.
In Proc. NeurIPS, 2023.
Patil et al. [2024]
S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez.
Gorilla: Large language model connected with massive APIs.
In Proc. NeurIPS, 2024.
Mistral AI [2026]
Mistral AI.
Mistral Vibe: A coding agent built on Mistral models.
https://github.com/mistralai/mistral-vibe (formerly mistralai/vibe), 2026.
Version 2.19.1, accessed July 2026.
Google [2026]
Google.
Gemini CLI: An open-source AI agent that brings the power of Gemini to the terminal.
https://github.com/google-gemini/gemini-cli, 2026.
Version 0.50.0, accessed July 2026; consumer service retired June 2026 in the announced transition to the closed-source Antigravity CLI.
Nous Research [2026]
Nous Research.
Hermes Agent: The agent that grows with you.
https://github.com/NousResearch/hermes-agent, 2026.
Version v2026.7.7.2 (0.18.2), accessed July 2026.
Zechner [2026]
M. Zechner.
Pi: The coding-agent harness you can make your own.
https://github.com/earendil-works/pi, 2026.
Version 0.80.6, accessed July 2026.
Anomaly [2026]
Anomaly (formerly SST).
OpenCode: The AI coding agent built for the terminal.
https://github.com/anomalyco/opencode (formerly sst/opencode), 2026.
Version 1.17.18, accessed July 2026.
Databricks [2026]
Databricks.
Introducing Omnigent: A meta-harness to combine, control and share your agents.
https://github.com/omnigent-ai/omnigent, June 2026.
Version 0.4.0, accessed July 2026.
xAI [2026]
xAI.
Grok Build: xAI’s coding agent CLI.
https://x.ai/cli, May 2026.
Accessed July 2026.
Claw Code Contributors [2026]
Claw Code Contributors.
Claw Code: A clean-room, provider-agnostic reimplementation of the Claude Code architecture.
https://claw-code.codes/, 2026.
Accessed July 2026.
Steinberger [2026]
P. Steinberger.
“You shouldn’t be prompting coding agents anymore. You should be designing loops that prompt your agents.”
Post on X, June 2026.
Macedo [2026]
S. O. de Macedo.
What makes a harness a harness: Necessary and sufficient conditions for an agent harness.
arXiv preprint arXiv:2606.10106, June 2026.
Rombaut [2026]
B. Rombaut.
Inside the scaffold: A source-code taxonomy of coding agent architectures.
arXiv preprint arXiv:2604.03515, April 2026.
Agent Skills [2025]
Agent Skills Working Group.
Agent Skills specification: file-system convention for capability bundles.
https://agentskills.io/, 2025.
Zhang et al. [2024a]
Y. Zhang, H. Ruan, Z. Fan, and A. Roychoudhury.
AutoCodeRover: Autonomous program improvement.
In Proc. ISSTA, 2024.
https://arxiv.org/abs/2404.05427.
Wang et al. [2025a]
X. Wang, B. Li, Y. Song, F. F. Xu, X. Tang, et al.
OpenHands: An open platform for AI software developers as generalist agents.
In Proc. ICLR, 2025.
https://arxiv.org/abs/2407.16741.
Chen et al. [2024a]
D. Chen, S. Lin, M. Zeng, D. Zan, et al.
CodeR: Issue resolving with multi-agent and task graphs.
arXiv preprint arXiv:2406.01304, 2024.
https://arxiv.org/abs/2406.01304.
Tao et al. [2024]
W. Tao, Y. Zhou, Y. Wang, W. Zhang, H. Zhang, and Y. Cheng.
MAGIS: LLM-based multi-agent framework for GitHub issue resolution.
In Proc. NeurIPS, 2024.
https://arxiv.org/abs/2403.17927.
Arora et al. [2024]
D. Arora, A. Sonwane, N. Wadhwa, A. Mehrotra, S. Utpala, R. Bairi, A. Kanade, and N. Natarajan.
MASAI: Modular architecture for software-engineering AI agents.
arXiv preprint arXiv:2406.11638, 2024.
https://arxiv.org/abs/2406.11638.
Ma et al. [2024]
Y. Ma, R. Cao, Y. Cao, Y. Zhang, J. Chen, Y. Liu, Y. Liu, B. Li, F. Huang, and Y. Li.
Lingma SWE-GPT: An open development-process-centric language model for automated software improvement.
arXiv preprint arXiv:2411.00622, 2024.
https://arxiv.org/abs/2411.00622.
Qian et al. [2024]
C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, et al.
ChatDev: Communicative agents for software development.
In Proc. ACL, 2024.
https://arxiv.org/abs/2307.07924.
Xia et al. [2024]
C. S. Xia, Y. Deng, S. Dunn, and L. Zhang.
Agentless: Demystifying LLM-based software engineering agents.
arXiv preprint arXiv:2407.01489, 2024.
https://arxiv.org/abs/2407.01489.
Antoniades et al. [2025]
A. Antoniades, A. Örwall, K. Zhang, Y. Xie, A. Goyal, and W. Wang.
SWE-Search: Enhancing software agents with Monte Carlo tree search and iterative refinement.
In Proc. ICLR, 2025.
https://arxiv.org/abs/2410.20285.
Zhang et al. [2024b]
K. Zhang, W. Yao, Z. Liu, Y. Feng, et al.
Diversity empowers intelligence: Integrating expertise of software engineering agents.
arXiv preprint arXiv:2408.07060, 2024.
https://arxiv.org/abs/2408.07060.
Shinn et al. [2023]
N. Shinn, F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, and S. Yao.
Reflexion: Language agents with verbal reinforcement learning.
In Proc. NeurIPS, 2023.
https://arxiv.org/abs/2303.11366.
Madaan et al. [2023]
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, et al.
Self-refine: Iterative refinement with self-feedback.
In Proc. NeurIPS, 2023.
https://arxiv.org/abs/2303.17651.
Gou et al. [2024]
Z. Gou, Z. Shao, Y. Gong, Y. Shen, Y. Yang, N. Duan, and W. Chen.
CRITIC: Large language models can self-correct with tool-interactive critiquing.
In Proc. ICLR, 2024.
https://arxiv.org/abs/2305.11738.
Chen et al. [2024b]
X. Chen, M. Lin, N. Schärli, and D. Zhou.
Teaching large language models to self-debug.
In Proc. ICLR, 2024.
https://arxiv.org/abs/2304.05128.
Qin et al. [2024]
Y. Qin, S. Liang, Y. Ye, K. Zhu, et al.
ToolLLM: Facilitating large language models to master 16000+ real-world APIs.
In Proc. ICLR, 2024.
https://arxiv.org/abs/2307.16789.
Du et al. [2024]
Y. Du, F. Wei, and H. Zhang.
AnyTool: Self-reflective, hierarchical agents for large-scale API calls.
In Proc. ICML, 2024.
https://arxiv.org/abs/2402.04253.
Zhang et al. [2023]
F. Zhang, B. Chen, Y. Zhang, J. Keung, J. Liu, D. Zan, Y. Mao, J.-G. Lou, and W. Chen.
RepoCoder: Repository-level code completion through iterative retrieval and generation.
In Proc. EMNLP, 2023.
https://arxiv.org/abs/2303.12570.
Liu et al. [2024a]
T. Liu, C. Xu, and J. McAuley.
RepoBench: Benchmarking repository-level code auto-completion systems.
In Proc. ICLR, 2024.
https://arxiv.org/abs/2306.03091.
Ding et al. [2023]
Y. Ding, Z. Wang, W. U. Ahmad, H. Ding, M. Tan, et al.
CrossCodeEval: A diverse and multilingual benchmark for cross-file code completion.
In Proc. NeurIPS Datasets and Benchmarks, 2023.
https://arxiv.org/abs/2310.11248.
Wang et al. [2025b]
Z. Z. Wang, A. Asai, X. V. Yu, F. F. Xu, Y. Xie, G. Neubig, and D. Fried.
CodeRAG-Bench: Can retrieval augment code generation?
In Findings of NAACL, 2025.
https://arxiv.org/abs/2406.14497.
Bogomolov et al. [2024]
E. Bogomolov, A. Eliseeva, T. Galimzyanov, E. Glukhov, et al.
Long Code Arena: A set of benchmarks for long-context code models.
arXiv preprint arXiv:2406.11612, 2024.
https://arxiv.org/abs/2406.11612.
Luo et al. [2024]
Q. Luo, Y. Ye, S. Liang, Z. Zhang, Y. Qin, et al.
RepoAgent: An LLM-powered open-source framework for repository-level code documentation generation.
In Proc. EMNLP System Demonstrations, 2024.
https://arxiv.org/abs/2402.16667.
Xi et al. [2023]
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, et al.
The rise and potential of large language model based agents: A survey.
arXiv preprint arXiv:2309.07864, 2023.
https://arxiv.org/abs/2309.07864.
Guo et al. [2024]
T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, and X. Zhang.
Large language model based multi-agents: A survey of progress and challenges.
In Proc. IJCAI Survey Track, 2024.
https://arxiv.org/abs/2402.01680.
Liu et al. [2025]
J. Liu, K. Wang, Y. Chen, X. Peng, Z. Chen, L. Zhang, and Y. Lou.
Large language model-based agents for software engineering: A survey.
ACM Trans. Softw. Eng. Methodol., 2025. To appear.
https://arxiv.org/abs/2409.02977.
Liu et al. [2024b]
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, et al.
AgentBench: Evaluating LLMs as agents.
In Proc. ICLR, 2024.
https://arxiv.org/abs/2308.03688.
Yao et al. [2023b]
S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan.
Tree of thoughts: Deliberate problem solving with large language models.
In Proc. NeurIPS, 2023.
https://arxiv.org/abs/2305.10601.
Wang et al. [2023]
L. Wang, W. Xu, Y. Lan, Z. Hu, Y. Lan, R. K.-W. Lee, and E.-P. Lim.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models.
In Proc. ACL, 2023.
https://arxiv.org/abs/2305.04091.
Zhang et al. [2024c]
Y. Zhang, R. Sun, Y. Chen, T. Pfister, R. Zhang, and S. Ö. Arik.
Chain of agents: Large language models collaborating on long-context tasks.
In Proc. NeurIPS, 2024.
https://arxiv.org/abs/2406.02818.
Zhou et al. [2024]
A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang.
Language agent tree search unifies reasoning, acting, and planning in language models.
In Proc. ICML, 2024.
https://arxiv.org/abs/2310.04406.
Packer et al. [2023]
C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez.
MemGPT: Towards LLMs as operating systems.
arXiv preprint arXiv:2310.08560, 2023.
https://arxiv.org/abs/2310.08560.
Jiang et al. [2023]
H. Jiang, Q. Wu, C.-Y. Lin, Y. Yang, and L. Qiu.
LLMLingua: Compressing prompts for accelerated inference of large language models.
In Proc. EMNLP, 2023.
https://arxiv.org/abs/2310.05736.
Debenedetti et al. [2024]
E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr.
AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents.
In Proc. NeurIPS Datasets and Benchmarks, 2024.
https://arxiv.org/abs/2406.13352.
Greshake et al. [2023]
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection.
In Proc. AISec, 2023.
https://arxiv.org/abs/2302.12173.
Yuan et al. [2024]
T. Yuan, Z. He, L. Dong, Y. Wang, et al.
R-Judge: Benchmarking safety risk awareness for LLM agents.
In Findings of EMNLP, 2024.
https://arxiv.org/abs/2401.10019.
Ye et al. [2024]
J. Ye, S. Li, G. Li, C. Huang, S. Gao, Y. Wu, Q. Zhang, T. Gui, and X. Huang.
ToolSword: Unveiling safety issues of large language models in tool learning across three stages.
In Proc. ACL, 2024.
https://arxiv.org/abs/2402.10753.
Mitra et al. [2024]
A. Mitra, L. Del Corro, G. Zheng, S. Mahajan, D. Rouhana, et al.
AgentInstruct: Toward generative teaching with agentic flows.
arXiv preprint arXiv:2407.03502, 2024.
https://arxiv.org/abs/2407.03502.
Gim et al. [2024]
I. Gim, G. Chen, S.-s. Lee, N. Sarda, A. Khandelwal, and L. Zhong.
Prompt cache: Modular attention reuse for low-latency inference.
In Proc. MLSys, 2024.
https://arxiv.org/abs/2311.04934.
Chen et al. [2024c]
L. Chen, M. Zaharia, and J. Zou.
FrugalGPT: How to use large language models while reducing cost and improving performance.
Trans. Mach. Learn. Res., 2024.
https://arxiv.org/abs/2305.05176.
Ong et al. [2025]
I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica.
RouteLLM: Learning to route LLMs with preference data.
In Proc. ICLR, 2025.
https://arxiv.org/abs/2406.18665.
Ding et al. [2024]
D. Ding, A. Mallick, C. Wang, R. Sim, S. Mukherjee, V. Rühle, L. V. S. Lakshmanan, and A. H. Awadallah.
Hybrid LLM: Cost-efficient and quality-aware query routing.
In Proc. ICLR, 2024.
https://arxiv.org/abs/2404.14618.
Lin et al. [2026]
X. Lin, C. Ruan, B. Rozière, M. Tufano, M. Velez, and B. Shen.
Agentic harness engineering: Observability-driven automatic evolution of coding-agent harnesses.
arXiv preprint arXiv:2604.25850, April 2026.
Guo et al. [2025]
H. Guo, Y. Hao, Y. Zhang, et al.
A measurement study of Model Context Protocol ecosystem.
arXiv preprint arXiv:2509.25292, 2025.
https://arxiv.org/abs/2509.25292.
Xu and Yan [2026]
R. Xu and Y. Yan.
Agent skills for large language models: Architecture, acquisition, security, and the path forward.
arXiv preprint arXiv:2602.12430, 2026.
https://arxiv.org/abs/2602.12430.
Saha and Hemanth [2026]
S. Saha and P. Hemanth.
Skilldex: A package manager and registry for agent skill packages with hierarchical scope-based distribution.
arXiv preprint arXiv:2604.16911, 2026.
https://arxiv.org/abs/2604.16911.
Guo et al. [2026]
Z. Guo, Z. Chen, et al.
SkillProbe: Security auditing for emerging agent skill marketplaces via multi-agent collaboration.
arXiv preprint arXiv:2603.21019, 2026.
https://arxiv.org/abs/2603.21019.
Kapoor et al. [2025]
S. Kapoor, N. Kolt, and S. Lazar.
Resist platform-controlled AI agents and champion user-centric agent advocates.
In Proc. ICML, 2025.
https://arxiv.org/abs/2505.04345.
Robbes et al. [2026]
R. Robbes, T. Matricon, et al.
Agentic much? Adoption of coding agents on GitHub.
arXiv preprint arXiv:2601.18341, 2026.
https://arxiv.org/abs/2601.18341.
Li et al. [2025b]
H. Li, H. Zhang, and A. E. Hassan.
The rise of AI teammates in Software Engineering 3.0: How autonomous coding agents are reshaping software engineering.
arXiv preprint arXiv:2507.15003, 2025.
https://arxiv.org/abs/2507.15003.