Large language models have improved automatic code generation, but generated programs still frequently contain syntax errors, incomplete edge-case handling, inefficient loops, and incorrect API usage. This study investigates reliable code generation and program repair through an iterative agent critique mechanism. We propose the Iterative Code Critique and Repair Model (ICCR), which contains four specialized agents: a task decomposition agent, a code generation agent, a test construction agent, and a repair optimization agent. The task decomposition agent converts natural-language requirements into input-output constraints and algorithmic steps. The code generation agent produces the initial program. The test construction agent creates unit tests, boundary cases, and mutation-based negative cases. The repair optimization agent revises the code according to compiler feedback, failed test traces, static analysis warnings, and complexity constraints. Experiments were conducted on 1,864 programming tasks from HumanEval, MBPP, APPS, and CodeContests, covering string processing, graph search, dynamic programming, numerical computation, data structures, and file-processing problems. A total of 18,640 generated programs and 92,300 automatically constructed test cases were evaluated. Compared with a single-agent GPT-style code generator, ICCR improved pass@1 from 63.8% to 78.6% on HumanEval and from 51.4% to 66.9% on MBPP. The compile-error rate decreased from 14.7% to 5.2%, while the average number of failed hidden test cases decreased by 38.5%. For algorithmic tasks in APPS, ICCR reduced average execution time by 21.6% and memory usage by 13.4% after optimization. Static analysis further showed a 17.9% reduction in cyclomatic complexity and a 24.3% reduction in duplicated code fragments. These results indicate that agent-level critique, test-driven repair, and complexity-aware optimization can improve both functional correctness and code quality in automated programming systems.
- Xiong, W., Guo, Y., & Zeng, Y. (2026). A Study on the Energy Efficiency of MEP Systems and Coordinated Dispatch Mechanisms for Multi-energy Systems in High-density Urban Buildings. Available at SSRN 7120518.
- Shivashankar, K., & Martini, A. (2025, November). Enhancing Python Code Maintainability Through Large Language Model-Based Approaches. In International Conference on Product-Focused Software Process Improvement (pp. 137-152). Cham: Springer Nature Switzerland.
- Hong, Z., Liang, S., Xu, T., & Chen, H. (2026). A Study on Failover Verification and Recovery Objective Prediction for Cross-Region Cloud Services.
- Bao, Y., Qiu, Y., & Wang, H. (2026). FLARE: A Real-Time Framework for Detecting Malicious Account Activities at Internet Scale. Available at SSRN 7185659.
- Zhang, Z., Wang, J., Li, Z., Wang, Y., & Zheng, J. (2025). Anncoder: A mti-agent-based code generation and optimization model. Symmetry, 17(7), 1087.
- Nimisha, P., & Eldho, K. J. (2025, September). Advancements in AI-Powered Code Generation: A Comparative Analysis of Modern LLMs. In 2025 4th International Conference on Innovative Mechanisms for Industry Applications (ICIMIA) (pp. 1492-1500). IEEE.
- Yuan, Y., Xu, T., Yin, J., & Huang, J. (2026). Modeling Loading Anomalies and Identifying Root Causes in Media Consumption Workflows on Large Social Media Platforms. Available at SSRN 7393318.
- Huang, J., Yin, J., Xu, T., & Yang, J. (2026). Quality Assessment and User Adoption Prediction for Enterprise-Level Financial Reporting Data Products. Available at SSRN 7179118.
- Özaytürk, H., & Buzluca, F. (2025, September). AutoStructor: A Generative AI-Based Framework for Automated Program Repair with Deep Learning-Guided Fault Localization. In 2025 10th International Conference on Computer Science and Engineering (UBMK) (pp. 1193-1198). IEEE.
- Zheng, J., & Makar, M. (2022). Causally motivated multi-shortcut identification and removal. Advances in Neural Information Processing Systems, 35, 12800-12812.
- Zhao, J., Fan, J., & Li, L. (2026). A Study on an Explainable Causal-Enhanced LLM Agent for Predicting the Forming Quality of Automotive Component Materials.
- You, S. (2026). Verifiable Audit Mechanisms in AI Compliance Automation: Scalability. Available at SSRN 6547458.
- Matias, B. C., Freire, S., Freitas, J., Fronchetti, F., Damevski, K., & Spinola, R. (2026). A Survey on Large Language Model Impact on Software Evolvability and Maintainability: the Good, the Bad, the Ugly, and the Remedy. arXiv preprint arXiv:2601.20879.
- Jiao, Y., Shi, T., Zhao, B., & Wang, A. (2026). Retrieval-Guided Structured Reasoning and Interpretable Representation Learning for Large-Scale Video–Language Models.
- Liang, S., Du, Y., Chen, W., & Liu, Z. (2026). Order Allocation and Emergency Transportation Decisions Amid Multi-Source Procurement Disruptions.
- Zhang, Z. (2026). A Study on the Impact of Cost Presentation on Decision-making Bias and Cancellation Behavior in Online Subscriptions. Available at SSRN 7349839.
- Varolia, J. (2026). Multi-Agent Collaborative Code Generator. Journal of Intelligent Decision Making and Information Science, 3(6s), 846-857.
- Wang, S., Feng, Y., & Fang, X. (2026, May). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving. In 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS) (pp. 253-256). IEEE.
- Yang, Q., & Du, Y. (2026). Predicting Rollback Risks and Controlling Gradual Rollout Traffic for Online Ranking Feature Deployments.
- Germano, L. B., Goldschmidt, R. R., Noya, R. C., & Duarte, J. C. (2025). A systematic review on detection, repair, and explanation of vulnerabilities in source code using large language models. IEEE Access, 13, 192263-192293.
- Zhao, Z., & Welsch, R. E. (2024). Hierarchical reinforced trader (hrt): A bi-level approach for optimizing stock selection and execution. arXiv preprint arXiv:2410.14927.
- Yuan, Y., Huang, J., Yin, J., & Xu, T. (2026). Confidence Calibration and Semantic Error Analysis for Ambiguous Coreference Resolution in User-generated Social Media Text. Available at SSRN 7393238.
- Siddiq, M. L., Ulfat, N., Raihan, N., Santos, J. C., & Zampieri, M. (2026). Multi-sallm: a multilingual security assessment of generated code. Automated Software Engineering, 33(4), 122.
- Qi, C., & Qiao, X. (2026). Building and Operating a Large Scale Multi-Agent System: A Case Study from Industry. Available at SSRN 6795198.
- Xu, T., Zhu, W., & Zhang, J. (2026). A Study on the Application of Alternative Data in Credit Assessment for the Unbanked Population.
- Gambacorta, L., Kwon, B., Park, T., Patelli, P., & Zhu, S. (2026). CB-LMs: language models for central banking. Journal of Financial Stability, 101585.
- Du, Y., Liu, Q., Dong, Y., & Chen, Z. (2026). Cross-Domain Transfer and Few-Shot Recognition of Defect Images in Complex Industrial Settings.
- Liang, S., Hsu, K. H., & Liu, Z. (2026). Rolling Scheduling and Capacity Balancing for Server Assembly Under Engineering Change Disturbances.
- Patcas, R., & Motogna, S. (2026). An evaluation study of large language models for addressing code quality issues. Empirical Software Engineering, 31(5), 118.
- Bao, Y., & Wang, H. (2026). Large-scale Metrics Migration: A Systematic Approach to Rebuilding Production Measurement Infrastructure. Available at SSRN 7185660.
- Journal
- Journal of Automated Discovery and Applied Informatics
- Volume
- 1 (2026)
- Issue
- 1 · Forthcoming issue
- Article number
- jadai20260007
- License
- CC BY 4.0