A build recipe is the set of things you need to build a project from scratch: the build system, the system and toolchain dependencies, and the phase commands and flags, in order. For a large share of real programs that recipe is not written down anywhere in the source tree. This project measures how much of it an LLM can recover from a cold start, and whether the difficulty of recovering it tracks how much each ecosystem declares.
- Type
- Independent research project
- Started
- May 2026 – ongoing
- Status
- Study design settled; harness in progress
- Build systems
- Autotools, Make, CMake, Maven, Gradle, Ant
- Evaluation set
- ~240 programs with hand-recorded recipes
- Corpora
- Apache, GNU, GitHub
- Python 3.12
- Docker
- Anthropic API
- Linux / WSL2
- Bash
The problem
Build reproducibility is an integral part of good software engineering but it is often difficult to achieve. Developers spend hours bug fixing and setting up projects, trying to work out whether their build failures are even their own fault in the first place [Huang-UnrelatedCI]. The instructions needed to build a project from scratch (which system packages to install, which toolchain version, which commands and in what order) often only exist in a working copy of the project, or in the head of whoever had it running before.
I want to be careful about how strong that claim is, because "most software is undocumented" is not something I can support. What I can support is a narrower version, measured on the build-systems dataset assembled in the lab I contributed to: a substantial fraction of the programs that do build require toolchain or dependency choices that are not recoverable from the files alone. The numbers behind that are in the hypothesis section.
Certain builds depend on information in the environment that the source tree does not have. A legacy Apache Java project can fail with a compilation error because it is only compatible with an older Java version, and nothing in the file tree mentions it. The only way a developer finds out is by building it, reading the error, and trying another Java version. That loop is what this project measures, not as a separate target, but as one failure category inside the scope below.
My approach
Given a project's source archive, the tool detects the build system, installs system dependencies, chooses a toolchain version, runs the build phases in order, and records the commands and flags needed to reproduce the result. Each recipe is verified by executing it in a fresh isolated container.
The design is two stages, and the split is the point. A deterministic script serves as the baseline, because that is what is most comparable to prior work [Agentless]. Where the script falls short is that it can only look up and reproduce recipes that already exist in the project files. From that point the LLM infers the build system, toolchain version, phase commands and flags by attempting the build and resolving its failures.
Stage A: deterministic baseline
Detect the build system by marker files, install the recorded system dependencies, run the default phase commands in a fresh container, and record success or failure per program. No model involved.
Stage B: LLM acting on failures
For each program the baseline cannot build, the model receives a summary of the source tree and the failure log from Stage A, and proposes a corrected recipe: flags, toolchain version, packages, or a fixed command. The system reruns in a fresh container and records whether it builds.
What the output actually is
This is the part I had wrong in an earlier draft. I was describing the output as "documentation", which is not something I can evaluate. There is no way to score whether a document lets someone "understand" a build. The output is a build recipe, as labelled fields, and it is scored two ways: does it build in a fresh container, and how does each field compare to the recipe a human recorded by hand for the same program.
A working Dockerfile is a replicable recipe, and I do not want to claim otherwise. The difference is that a Dockerfile collapses everything into one artifact. Separating the build system, system dependencies, toolchain version and per-phase commands into named fields is what makes per-component scoring possible: build system classification accuracy, dependency precision and recall, toolchain version match, and per-phase outcomes, rather than a single build-succeeded bit.
The output format the harness is being built to produce, on an illustrative program:
{
"name": "accumulo",
"corpus": "Apache",
"detected_build_system": "Maven",
"apt_dependencies": ["maven", "openjdk-17-jdk"],
"tool_versions": { "java": "openjdk 17.0.13", "maven": "3.8.7" },
"commands": {
"bootstrap": "mvn process-resources",
"configure": null,
"compile": "mvn process-classes",
"test": "mvn test"
},
"resolution": {
"solved_by": "llm",
"baseline_result": "fail_at_compile",
"attempts": [
{ "stage": "A", "toolchain": {"java": "openjdk-21-jdk"},
"failed_phase": "compile",
"error_signature": "warnings found and -Werror specified" },
{ "stage": "B", "fix": {"field": "toolchain.java", "to": "openjdk-17-jdk"},
"rationale": "JDK 21 errors on deprecated-for-removal API; targets JDK 17",
"result": "success" }
]
},
"ground_truth_match": {
"build_system": true,
"apt_dependencies": { "precision": 1.0, "recall": 1.0 },
"toolchain_java": true
}
}
The hypothesis this maps
Porting a technique from one ecosystem to another is engineering, not research, unless there is something about the other ecosystems that makes the problem intrinsically different. I think there is, and it is the hypothesis the study is built around:
Cold-start inference degrades predictably with how much dependency information an ecosystem declares. Maven and Gradle declare their dependencies, so a model can read them. Autotools and Make declare almost nothing, so dependencies only surface as configure or compile errors. Inference should be structurally easier where dependencies are stated and harder where they are latent, and the contribution is mapping that gradient and explaining why.
That also predicts which fields are easy. A build system is declared by the
presence of pom.xml or configure.ac, so it should be classified
nearly perfectly. A required Java Development Kit (JDK) version is often declared nowhere and only surfaces as a
compile error, so it should be recovered poorly by anything that does not actually attempt
the build and read the error.
The evidence that the non-scriptable part of this process is large, drawn from the lab dataset:
- Of the 161 Apache projects with a recorded required Java version, 92 require JDK 8, 11 or 17 instead of the default JDK 21 (JDK 8: 57, JDK 11: 17, JDK 17: 18).
- 103 programs where the default build commands fail outright.
- 19 programs that do not compile under the default
gcc/g++flags. - 25 of the included Apache projects do not succeed with just
mvn verify. -
Configure invocations that depend on less obvious compatibility flags such as
CFLAGS+=-fcommon ./configureandFORCE_UNSAFE_CONFIGURE=1 ./configure.
None of these corrections are recoverable from filenames alone, and none would survive being collapsed into a single Dockerfile.
Prior work, and what I am measured against
Repo2Run builds executable Docker environments for Python repositories, reporting an 86% success rate on 420 repositories [Repo2Run]. CXXCrafter targets C and C++ projects specifically and reports a 78% build success rate [CXXCrafter]. These are the two systems closest to mine, and the comparison design runs my tool and theirs over each other's program sets (mine over their corpora, theirs over the apt-installable subset) so that a reported success rate can be read against a corpus other than the one its own authors chose.
Lyu et al. is the nearest work overall: an LLM generates a working Dockerfile by iteratively patching until the project builds, then optimises image size [Lyu-Dockerfile]. Related work from the same authors fixes missing dependency declarations in existing build files [Lyu-MDfixer], and Rosa et al. trained a transformer to generate Dockerfiles from natural-language requirement specifications [Rosa]. Three differences keep this project distinct:
- Scale and distribution. They evaluate on 35 projects across five languages at roughly 80% build success. The apt-installable subset here is around 240 programs, with 109 Maven and 184 Autotools/Make.
- What is measured. They measure buildability, image size and build time. I measure the inferred recipe against a human-recorded one, per component.
- What the model is given. Working from the raw source tree rather than from a README or other natural-language artifact is a further difference, one I still need to confirm against their full paper rather than the abstract.
GradleFixer repairs Android Gradle build failures through domain-specific tools instead of a general-purpose shell, reporting an 81.4% resolve rate on AndroidBuildBench and outperforming a shell-based agent [GradleFixer]. That is where the tool-wrapping part of this design comes from, and testing whether it holds across more than a single ecosystem is one of the research questions.
Two more shape the design without being direct comparisons. Agentless implements a fixed localization–repair–validation pipeline and stays competitive with more complex agents on SWE-bench Lite [Agentless]; that is the logic the deterministic baseline follows. ESAA keeps an append-only log of validated intents for auditability and deterministic replay [ESAA], which is what inspired the recorder here. EChecker [EChecker] and Breaking-Good [Breaking-Good] detect and explain build problems without repairing the environment; they are adjacent rather than competing, and I no longer lean on them to make the case for this project.
The work that started this is Foreman [Foreman], which models per-phase file and behaviour permissions to harden build systems against pipeline poisoning. Its specifications are currently hand-written and automated specification inference is listed as future work. A successful build trace from this project could help draft those specifications, but that is a stretch goal and out of scope for now.
Research questions
- RQ1. Given only a source archive, how effectively can an LLM infer a build recipe that produces a successful build in an isolated container, and how does that vary across ecosystems and corpora?
- RQ2. Measured on success rate, robustness and LLM-call cost, does letting the model retry on failure outperform a single repair attempt, or does the deterministic baseline accomplish this alone?
- RQ3. Does constraining the model to domain-specific tool wrappers, instead of raw shell access, improve build success and reduce invalid or unsafe actions, and does that hold beyond GradleFixer's single Android/Gradle setting?
- RQ4. For programs neither the baseline nor the model can build, which failure categories are reachable with more attempts or different toolchain choices, and which are genuinely out of scope?
Limitations
Stating the limits before being asked is the point, so here they are up front.
- The evaluation set only includes apt-installable programs, which removes the harder dependency cases. Programs needing pip, npm or cargo, specific hardware, or a legacy version that is no longer packaged are out of bounds, and 141 of the 1002 programs in the dataset are excluded on those grounds. Within this scope the study can still measure detection accuracy, toolchain version selection, phase command correctness, and recovery through flag and version changes.
- The build systems are heavily skewed. Of the 368 programs that build successfully with full recorded recipes: Autotools and Make (184), Maven (109), CMake (36), Make (29), Gradle (6), Ant (4). Gradle and Ant results should be read as suggestive rather than statistically solid.
- The human baseline was produced by humans. It is subject to human error, and the recorded recipes are not guaranteed to be the only correct ones, or the most efficient ones.
- The cross-tool comparison depends on the other systems being runnable. Until their artifacts are confirmed available this remains a comparison design, not a completed result.
- The ground-truth dataset is one I contributed to, not one I built alone. It came out of lab work with multiple contributors.
Status
Study design settled; environment and harness in progress; results pending. Splitting what is designed from what is running is deliberate. It is the same honesty that makes the evaluation argument credible in the first place.
Settled now
- Environment and repository setup: Windows 11 host on Windows Subsystem for Linux 2 (WSL2) with Ubuntu 24.04, Docker driven from Python through the Docker software development kit (SDK), and the dataset subset stored as JavaScript Object Notation (JSON).
- Six build systems worked with directly, hands-on.
- Build-failure diagnosis and failure-category construction.
- Study design, metric definition, and the two-stage evaluation plan.
- Literature synthesis across the existing systems in this space.
In progress
- The containerized executor and build-system detection.
- The LLM inference call and failure-recovery loop with attempt budgets.
- Per-ecosystem results, the tool-wrapper ablation, and the finished map.
Still open
- Whether Lyu et al. rely on a README or other natural-language artifact when generating a Dockerfile, or work from the raw repository as this project does.
- Whether to include a smaller, harder non-apt subset now, so that the dependency boundary is visible rather than assumed.
- Whether a multi-iteration loop with pass@k and robustness curves is worth building, or whether it only earns its keep if the single attempt shows promise first.
What changed after review
An earlier version of this proposal led with documentation and information loss. That framing did not survive review, and the objections were the right ones:
Documentation → build recipe
Documentation is hard to evaluate; there is no defensible way to score whether a document lets someone understand a build. The output is now a recipe with labelled fields, scored on execution in a fresh container and on per-field agreement with the human record.
Rejected: evaluating comprehension or documentation quality.
Multi-ecosystem coverage → the declaration gradient
"It works on more ecosystems" is porting, not a research contribution. The contribution is the claim that inference difficulty is structurally tied to how much an ecosystem declares, and a map that shows where that holds.
Rejected: breadth of ecosystem coverage as the headline claim.
An unsupported premise → dataset numbers
"Most software lacks appropriate documentation" was asserted without evidence. It is now a narrower, measured claim about the programs in the dataset, with the counts stated above.
Rejected: the broad claim about software in general.
References
- [Agentless] Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, Lingming Zhang. "Agentless: Demystifying LLM-based Software Engineering Agents." arXiv:2407.01489, 2024.
- [GradleFixer] Ha Min Son, Huan Ren, Xin Liu, Zhe Zhao. "Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools." arXiv:2510.08640, 2025.
- [Lyu-Dockerfile] Jun Lyu, He Zhang, Yusong Yuan, Lanxin Yang, Yue Li, Manuel Rigger. "Automatic Dockerfile Generation with Large Language Models." ICSE 2026.
- [Lyu-MDfixer] Jun Lyu, He Zhang, Lanxin Yang, Yue Li, Chenxing Zhong, Manuel Rigger. "Automatic Fixing of Missing Dependency Errors." ASE 2025. DOI:10.1109/ASE63991.2025.00056.
- [EChecker] Jun Lyu, Shanshan Li, He Zhang, Yang Zhang, Guoping Rong, Manuel Rigger. "Detecting Build Dependency Errors in Incremental Builds." ISSTA 2024. DOI:10.1145/3650212.3652105. Also arXiv:2404.13295.
- [Breaking-Good] Frank Reyes, Benoit Baudry, Martin Monperrus. "Breaking-Good: Explaining Breaking Dependency Updates with Build Analysis." SCAM 2024, pp. 36–46. DOI:10.1109/SCAM63643.2024.00014. Also arXiv:2407.03880.
- [Rosa] Giovanni Rosa, Antonio Mastropaolo, Simone Scalabrino, Gabriele Bavota, Rocco Oliveto. "Automatically Generating Dockerfiles via Deep Learning: Challenges and Promises." ICSSP 2023, pp. 1–12. DOI:10.1109/ICSSP59042.2023.00011. Also arXiv:2303.15990.
- [Repo2Run] Ruida Hu, Chao Peng, Xinchen Wang, Junjielong Xu, Cuiyun Gao. "Repo2Run: Automated Building Executable Environment for Code Repository at Scale." NeurIPS 2025. Also arXiv:2502.13681.
- [CXXCrafter] Zhengmin Yu, Yuan Zhang, Ming Wen, Yinan Nie, Wenhui Zhang, Min Yang. "CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building." Proc. ACM Softw. Eng. 2 (FSE), pp. 2618–2640, 2025. DOI:10.1145/3729386.
- [ESAA] Elzo Brito dos Santos Filho. "ESAA: Event Sourcing for Autonomous Agents in LLM-Based Software Engineering." arXiv:2602.23193, 2026.
- [Foreman] Brent Pappas, Paul Gazzillo. "Build Code is Still Code: Finding the Antidote for Pipeline Poisoning." ICSE 2026. Also arXiv:2601.08995.
- [Huang-UnrelatedCI] Andie Huang, Daniel Alencar da Costa, Grant Dick, Mariam El Mezouar. "Is this Build Failure Related to my Patch? An Empirical Study of Unrelated Build Failures in Continuous Integration." arXiv:2605.05564, 2026.
More work on the projects page, or back to Selected Work.