PaperSpread Research Publishing
Journal of Automated Discovery and Applied Informatics

Reproducible Scientific Workflow Generation through Agentic Code Verification

Read & download PDF
Abstract

Scientific data analysis increasingly depends on code, but manually written analysis scripts often suffer from missing preprocessing steps, inconsistent statistical tests, undocumented parameters, and poor reproducibility. Large language models can generate analysis code from natural-language research questions, yet the generated workflows may execute incorrectly or produce results that cannot be reproduced across environments. This study explores reproducible scientific workflow generation through agentic code verification. We propose SciFlow-Agent, a multi-agent model containing a data inspection agent, a statistical planning agent, a workflow code generation agent, a visualization agent, and a reproducibility verification agent. The data inspection agent identifies variable types, missing values, outliers, and table relationships. The statistical planning agent selects suitable methods, including regression, survival analysis, time-series modeling, hypothesis testing, and non-parametric comparison. The workflow code generation agent produces executable Python and R scripts. The visualization agent generates publication-ready figures, while the verification agent reruns the workflow under controlled environments and checks numerical consistency, dependency completeness, and result traceability. Experiments were conducted on 740 public datasets from biomedical research, environmental science, public health, economics, and social science. The benchmark included 2,960 analysis tasks, 8,420 expected statistical outputs, and 5,180 figure-generation requirements. Compared with a single-agent code generation model, SciFlow-Agent improved full-workflow execution success from 57.3% to 79.5%. Reproducible result matching increased from 61.8% to 84.2%, and statistical-method selection errors decreased by 36.7%. The average absolute deviation between regenerated numerical results and reference outputs was reduced from 0.084 to 0.031. Dependency-related failures decreased by 48.6%, and figure specification accuracy improved from 70.4% to 86.9%. These results suggest that agentic code verification can strengthen the reliability, transparency, and reproducibility of automated scientific programming.

Keywords
Scientific workflowcode generationreproducibilitystatistical computingautomated data analysiscode verificationmulti-agent systems
References
  1. Ahmad, R., Manne, N. N., & Malik, T. (2025). Improving reproducibility of interactive notebooks using application virtualization. Future Generation Computer Systems, 108043.
  2. Zhao, J., Fan, J., & Li, L. (2026). A Study on an Explainable Causal-Enhanced LLM Agent for Predicting the Forming Quality of Automotive Component Materials.
  3. You, S. (2026). Verifiable Audit Mechanisms in AI Compliance Automation: Scalability. Available at SSRN 6547458.
  4. Yao, G., Zheng, J., Wang, Z., Zhang, W., Han, R., Zhao, C., ... & Liu, R. (2026, March). V-pruner: A fast and globally-informed token pruning framework for vision transformer. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, No. 40, pp. 34396-34404).
  5. Collyer, T. A. (2026). Epistemic cultures, knowledge types, and disciplinary tensions in population health research: An international interview study. Social Science & Medicine, 119786.
  6. Jiao, Y., Shi, T., Zhao, B., & Wang, A. (2026). Retrieval-Guided Structured Reasoning and Interpretable Representation Learning for Large-Scale Video–Language Models.
  7. Liang, S., Du, Y., Chen, W., & Liu, Z. (2026). Order Allocation and Emergency Transportation Decisions Amid Multi-Source Procurement Disruptions.
  8. Zhang, Z. (2026). A Study on the Impact of Cost Presentation on Decision-making Bias and Cancellation Behavior in Online Subscriptions. Available at SSRN 7349839.
  9. Gülmez, B. (2026). Code generation with large language models: a survey from neural program synthesis to autonomous software development. Applied Intelligence, 56(6), 200.
  10. Qi, C., & Qiao, X. (2026). Building and Operating a Large Scale Multi-Agent System: A Case Study from Industry. Available at SSRN 6795198.
  11. Du, Y., Liu, Q., Dong, Y., & Chen, Z. (2026). Cross-Domain Transfer and Few-Shot Recognition of Defect Images in Complex Industrial Settings.
  12. Naseem, R., & Zaki, M. (2026). Automated Code Generation Using Large Language Models: A Comparative Study. Research Consortium Archive, 4(2), 1178-1186.
  13. Xiong, W., Guo, Y., & Zeng, Y. (2026). A Study on the Energy Efficiency of MEP Systems and Coordinated Dispatch Mechanisms for Multi-energy Systems in High-density Urban Buildings. Available at SSRN 7120518.
  14. Liang, S., Zhang, X., Du, Y., & Chen, W. (2026). Supply Network Restructuring and Capacity Ramp-Up Limits in the Localized Expansion of High-Performance Computing Equipment Production. Available at SSRN 7339738.
  15. Bao, Y., & Wang, H. (2026). Large-scale Metrics Migration: A Systematic Approach to Rebuilding Production Measurement Infrastructure. Available at SSRN 7185660.
  16. Chopra, S., Basra, E. N. K., Simran, E., & Mohapatra, E. D. (Eds.). (2025). Fundamentals of Data Handling and Visualization. City, 100(15), 00-000.
  17. Zhao, Z., & Welsch, R. E. (2024). Hierarchical reinforced trader (hrt): A bi-level approach for optimizing stock selection and execution. arXiv preprint arXiv:2410.14927.
  18. Yuan, Y., Xu, T., Yin, J., & Huang, J. (2026). Modeling Loading Anomalies and Identifying Root Causes in Media Consumption Workflows on Large Social Media Platforms. Available at SSRN 7393318.
  19. Bao, Y., & Wang, H. (2026). Orion: A High-Throughput, Fault-Tolerant Event Routing Architecture for Hyperscale Data Streams. Fault-Tolerant Event Routing Architecture for Hyperscale Data Streams (January 01, 2026).
  20. Otoum, N., & Elkhalili, N. (2026). Methods and techniques of agentic software engineering: A systematic literature review. IEEE Access, 14, 7443-7465.
  21. Wang, S., Feng, Y., & Fang, X. (2026, May). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving. In 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS) (pp. 253-256). IEEE.
  22. Yang, Q., & Du, Y. (2026). Predicting Rollback Risks and Controlling Gradual Rollout Traffic for Online Ranking Feature Deployments.
  23. Ferdaus, M. M., Abdelguerfi, M., Loup, E., N. Niles, K., Pathak, K., & Sloan, S. (2026). Towards trustworthy ai: A review of ethical and robust large language models. ACM Computing Surveys, 58(7), 1-43.
  24. Yuan, Y., Yin, J., Huang, J., & Xu, T. (2026). Weakly Supervised Anomaly Detection and Privacy Risk Scoring in Community Message Threads.
  25. Hong, Z., Liang, S., Xu, T., & Chen, H. (2026). A Study on Failover Verification and Recovery Objective Prediction for Cross-Region Cloud Services.
  26. Trofimova, E., Shamina, Z., Selifanova, M., Zaitsev, A., Savchuk, R., Minets, M., ... & Ustyuzhanin, A. E. (2025). ML2B: Multi-Lingual ML Benchmark For AutoML. arXiv preprint arXiv:2509.22768.
  27. Wu, J., Wu, D., Zhang, J., & Peng, Y. (2026). Generative AI Feedback in Junior High School Artistic Creation Evaluation. Available at SSRN 7244838.
  28. Huang, J., Xu, T., Yin, J., & Yang, J. (2026). Data Quality Monitoring and Intelligent Scheduling Optimization in Financial Data Workflow Automation. Available at SSRN 7179338.
  29. Souaifi, M., Dhahbi, W., Jebabli, N., Ceylan, H. İ., Boujabli, M., Muntean, R. I., & Dergaa, I. (2025). Artificial intelligence in sports biomechanics: A scoping review on wearable technology, motion analysis, and injury prevention. Bioengineering, 12(8), 887.
  30. Zhang, Z., Wang, J., Li, Z., Wang, Y., & Zheng, J. (2025). Anncoder: A mti-agent-based code generation and optimization model. Symmetry, 17(7), 1087.
Publication details
Journal
Journal of Automated Discovery and Applied Informatics
Volume
1 (2026)
Issue
1 · Forthcoming issue
Article number
jadai20260008
License
CC BY 4.0