Abstract
In this paper, we develop a multi-agent pipeline-based approach for solving competitive programming problems via Large Language Models (LLMs). Specifically, we analyze the Writer-Reviewer pipeline where the Writer Agent produces Python solutions and the Reviewer Agent gives static natural language feedback. Experiments on a 96-problem AtCoder subset of LiveCodeBench involve comparing twelve pipeline setups which employ three different models (GPT-OSS-20B, Qwen3-Coder-30B-A3B-Instruct, and Qwen2.5-Coder-7B-Instruct) in various agent roles. Within this experimental setup, the review process increases the performance of GPT-OSS-20B from 87.5% to 91.7% Pass@1, with the bootstrap 95% confidence intervals overlapping, while the performance of Qwen3-Coder-30B-A3B-Instruct does not improve and the performance of Qwen2.5-Coder-7B-Instruct improves marginally from 0% to 1.0%. This may indicate that the static iterative feedback mechanism helps to further improve the performance of a strong Writer but not the performance of a weak Writer. A diagnostic audit suggests that the very low Qwen2.5-Coder-7B-Instruct scores mainly reflect structured-output compliance failures in our setup rather than standalone coding capability. Moreover, performance varies more when changing the Writer Agent than when changing the Reviewer Agent, with differences of up to 88.6 percentage points across Writers and up to 17.7 percentage points across Reviewers.
References
Anthropic, 2024 The Claude 3 Model Family: A New Standard for Intelligence. Technical report, Anthropic.
Austin, J., A. Odena, M. Nye, M. Bosma, H. Michalewski, et al., 2021 Program synthesis with large language models.
Azerbayev, Z., H. Schoelkopf, K. Paster, M. D. Santos, S. McAleer, et al., 2024 Llemma: An open language model for mathematics.
Birhane, A., A. Kasirzadeh, D. Leslie, and S. Wachter, 2023 Science in the age of large language models. Nature Reviews Physics 5.
Brown, T. B., B. Mann, N. Ryder, M. Subbiah, J. Kaplan, et al., 2020 Language models are few-shot learners.
Chen, M., J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, et al., 2021 Evaluating large language models trained on code.
Chen, X., M. Lin, N. Schärli, and D. Zhou, 2023 Teaching large language models to self-debug.
Cui, J., M. Ning, Z. Li, B. Chen, Y. Yan, et al., 2024 ChatLaw: A multi-agent collaborative legal assistant with knowledge graph enhanced mixture-of-experts large language model.
DeepSeek-AI, 2025 DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.
Gemini Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, et al., 2023 Gemini: A Family of Highly Capable Multimodal Models.
Guo, D., Q. Zhu, D. Yang, Z. Xie, K. Dong, et al., 2024 DeepSeek-Coder: When the large language model meets programming – the rise of code intelligence.
Hendrycks, D., S. Basart, S. Kadavath, M. Mazeika, A. Arora, et al., 2021 Measuring coding challenge competence with APPS.
Hong, S., M. Zhuge, J. Chen, X. Zheng, Y. Cheng, et al., 2024 MetaGPT: Meta programming for a multi-agent collaborative framework.
Huang, D., J. M. Zhang, M. Luck, Q. Bu, Y. Qing, et al., 2024a AgentCoder: Multi-agent-based code generation with iterative testing and optimisation.
Huang, J., X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, et al., 2024b Large language models cannot self-correct reasoning yet.
Hui, B., J. Yang, Z. Cui, J. Yang, D. Liu, et al., 2024 Qwen2.5-Coder Technical Report.
Ishibashi, Y. and Y. Nishimura, 2024 Self-organized agents: A LLM multi-agent framework toward ultra large-scale code generation and optimization.
Islam, M. A., M. E. Ali, and M. R. Parvez, 2024 MapCoder: Multi-agent code generation for competitive problem solving.
Jain, N., K. Han, A. Gu, W.-D. Li, F. Yan, et al., 2024 LiveCodeBench: Holistic and contamination-free evaluation of large language models for code.
Jin, H., Z. Sun, and H. Chen, 2024 RGD: Multi-LLM based agent debugger via refinement and generation guidance.
Kasneci, E., K. Sessler, S. Küchemann, M. Bannert, D. Dementieva, et al., 2023 ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences 103: 102274.
Labrak, Y., A. Bazoge, E. Morin, P.-A. Gourraud, M. Rouvier, et al., 2024 BioMistral: A collection of open-source pretrained large language models for medical domains.
LangChain, 2025 LangGraph Documentation. Online, Accessed: 2026-04-30.
Li, R., L. B. Allal, Y. Zi, N. Muennighoff, D. Kocetkov, et al., 2023 StarCoder: May the source be with you!
Li, Y., D. Choi, J. Chung, N. Kushman, J. Schrittwieser, et al., 2022 Competition-level code generation with AlphaCode. Science 378: 1092–1097.
Lozhkov, A., R. Li, L. B. Allal, F. Cassano, J. Lamy-Poirier, et al., 2024 StarCoder 2 and The Stack v2: The next generation. arXiv preprint arXiv:2402.19173.
Madaan, A., N. Tandon, P. Gupta, S. Hallinan, L. Gao, et al., 2023 Self-Refine: Iterative refinement with self-feedback.
OpenAI, 2024 Learning to Reason with LLMs. Accessed: 2026-04-30.
OpenAI, 2025 GPT-OSS-120B & GPT-OSS-20B Model Card. Includes GPT-OSS-20B parameter counts and June 2024 knowledge cutoff.
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, et al., 2024 GPT-4 Technical Report.
Qian, C., W. Liu, H. Liu, N. Chen, Y. Dang, et al., 2024 ChatDev: Communicative agents for software development.
Quan, S., J. Yang, B. Yu, B. Zheng, D. Liu, et al., 2025 CodeELO: Benchmarking competition-level code generation of LLMs with human-comparable Elo ratings.
Qwen Team, 2024 Qwen2.5-Coder-7B-Instruct Model Card. Online, Accessed: 2026-04-30.
Qwen Team, 2025 Qwen3-Coder-30B-A3B-Instruct Model Card. Online, Accessed: 2026-04-30.
Rozière, B., J. Gehring, F. Gloeckle, S. Sootla, I. Gat, et al., 2024 Code Llama: Open foundation models for code.
Shao, Z., P. Wang, Q. Zhu, R. Xu, J. Song, et al., 2024 DeepSeekMath: Pushing the limits of mathematical reasoning in open language models.
Shinn, N., F. Cassano, E. Berman, A. Gopinath, K. Narasimhan, et al., 2023 Reflexion: Language agents with verbal reinforcement learning.
Singhal, K., S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, et al., 2022 Large language models encode clinical knowledge.
Stack Overflow, 2024 Stack Overflow Developer Survey 2024. Online, Accessed: 2026-04-30.
Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, et al., 2017 Attention is all you need. In Advances in Neural Information Processing Systems, volume 30.
Zhong, L., Z. Wang, and J. Shang, 2024 Debug like a human: A large language model debugger via verifying runtime execution step-by-step.

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
