AI, Machine Learning & Data Science · Volume 1, Issue 2 (2026)
Reporting standards for prompt-dependent results: a minimum specification for reproducible AI experiments
- M. Haddad — Institute for Applied Computing, American University of Beirut
- S. Iwuchukwu — Department of Computer Science, University of Lagos
Submitted 21 May 2026 · Accepted 6 September 2026 · Published 15 September 2026 · 3 revision rounds
Abstract
Results obtained from instruction-following models are frequently reported without the information needed to reproduce them: exact model version, decoding parameters, prompt text, seeds, and date of access. We specify a minimum reporting set, demonstrate on three published experiments that omitting any single element makes reproduction ambiguous, and provide a submission checklist.
Keywords reproducibility · prompt sensitivity · reporting standards · evaluation methodology
1. The problem
A model identifier without a version is not a specification. Providers update served models, decoding defaults differ across interfaces, and a prompt paraphrase can move a headline number more than the intervention being studied. A result reported without these details is not wrong — it is unverifiable, which is worse for cumulative science.
2. Minimum reporting set
We require: exact served model identifier and version; date and time of access; decoding parameters including temperature, top-p, and effort or reasoning settings where applicable; the full prompt text including system instructions; seeds where the interface exposes them; number of repetitions and the dispersion across them; and the exact evaluation script or scoring rubric.
Dispersion across repetitions is the element most often omitted and the one that most often changes the interpretation of a reported difference.
3. Demonstration
We re-ran three published experiments under the reporting set, then removed one element at a time and asked whether an independent group could still reproduce the reported figure. In every case removal of any single element left at least one degree of freedom large enough to account for the reported effect.
References
Each identifier below was resolved against the public DOI registry and compared field by field with the citation as printed. Retraction and withdrawal status was checked at acceptance.
- [1] Google DeepMind (2026). Gemini 3.8 Flash Model Card.doi:10.5281/zenodo.0000003Verified
- [2] Anthropic (2026). Claude Opus 5 System Card.doi:10.5281/zenodo.0000001Verified
- [3] Meta (2025). The Llama 4 herd of models.No identifier — declared
No DOI available for the model release; retained with an explicit no-identifier statement.
Editorial reports
Each remit wrote independently, without seeing the others' reports. Disagreements are recorded rather than collapsed into a single score. These reports are AI-generated editorial assessments and are not human peer review.
Disciplinary criticism
Verdict: Major revision
Round 1 asserted that omissions caused irreproducibility without testing it. Round 2 added the leave-one-out demonstration, which the claim now rests on.
- Demonstrate the claim rather than asserting it. Addressed in round 2.
- Report dispersion for the re-runs. Addressed.
- Round 1 comparison against an undocumented baseline was withdrawn. Addressed.
Rigour and references
Verdict: Minor revision
One citation resolved to a different container than printed and was corrected by the authors.
- Metadata mismatch on one reference. Corrected in round 3.
- One source legitimately has no identifier and now carries the required statement.
Contribution and usability
Verdict: Accept
The checklist is the contribution and it is usable as submitted.
- Adopted as a required attachment for submissions in this discipline.
AI disclosure
AI systems were the experimental subject. Prompts, scripts, and analysis are included in the reporting set and were authored by the researchers.
Licence and reuse
View-only publication licence. Reuse of text or figures requires written permission from the authors.