Mapping How LLMs Build Software

An applicability study: where large language model (LLM) build-recipe inference works, where it fails, and whether the difference tracks how much each ecosystem declares about its own dependencies.

A build recipe is the set of things you need to build a project from scratch: the build system, the system and toolchain dependencies, and the phase commands and flags, in order. For a large share of real programs that recipe is not written down anywhere in the source tree. This project measures how much of it an LLM can recover from a cold start, and whether the difficulty of recovering it tracks how much each ecosystem declares.

Type
Independent research project
Started
May 2026 – ongoing
Status
Study design settled; harness in progress
Build systems
Autotools, Make, CMake, Maven, Gradle, Ant
Evaluation set
~240 programs with hand-recorded recipes
Corpora
Apache, GNU, GitHub

The problem

Build reproducibility is an integral part of good software engineering but it is often difficult to achieve. Developers spend hours bug fixing and setting up projects, trying to work out whether their build failures are even their own fault in the first place [Huang-UnrelatedCI]. The instructions needed to build a project from scratch (which system packages to install, which toolchain version, which commands and in what order) often only exist in a working copy of the project, or in the head of whoever had it running before.

I want to be careful about how strong that claim is, because "most software is undocumented" is not something I can support. What I can support is a narrower version, measured on the build-systems dataset assembled in the lab I contributed to: a substantial fraction of the programs that do build require toolchain or dependency choices that are not recoverable from the files alone. The numbers behind that are in the hypothesis section.

Certain builds depend on information in the environment that the source tree does not have. A legacy Apache Java project can fail with a compilation error because it is only compatible with an older Java version, and nothing in the file tree mentions it. The only way a developer finds out is by building it, reading the error, and trying another Java version. That loop is what this project measures, not as a separate target, but as one failure category inside the scope below.

My approach

Given a project's source archive, the tool detects the build system, installs system dependencies, chooses a toolchain version, runs the build phases in order, and records the commands and flags needed to reproduce the result. Each recipe is verified by executing it in a fresh isolated container.

The design is two stages, and the split is the point. A deterministic script serves as the baseline, because that is what is most comparable to prior work [Agentless]. Where the script falls short is that it can only look up and reproduce recipes that already exist in the project files. From that point the LLM infers the build system, toolchain version, phase commands and flags by attempting the build and resolving its failures.

Stage A: deterministic baseline

Detect the build system by marker files, install the recorded system dependencies, run the default phase commands in a fresh container, and record success or failure per program. No model involved.

Stage B: LLM acting on failures

For each program the baseline cannot build, the model receives a summary of the source tree and the failure log from Stage A, and proposes a corrected recipe: flags, toolchain version, packages, or a fixed command. The system reruns in a fresh container and records whether it builds.

What the output actually is

This is the part I had wrong in an earlier draft. I was describing the output as "documentation", which is not something I can evaluate. There is no way to score whether a document lets someone "understand" a build. The output is a build recipe, as labelled fields, and it is scored two ways: does it build in a fresh container, and how does each field compare to the recipe a human recorded by hand for the same program.

A working Dockerfile is a replicable recipe, and I do not want to claim otherwise. The difference is that a Dockerfile collapses everything into one artifact. Separating the build system, system dependencies, toolchain version and per-phase commands into named fields is what makes per-component scoring possible: build system classification accuracy, dependency precision and recall, toolchain version match, and per-phase outcomes, rather than a single build-succeeded bit.

The output format the harness is being built to produce, on an illustrative program:

{
  "name": "accumulo",
  "corpus": "Apache",
  "detected_build_system": "Maven",
  "apt_dependencies": ["maven", "openjdk-17-jdk"],
  "tool_versions": { "java": "openjdk 17.0.13", "maven": "3.8.7" },
  "commands": {
    "bootstrap": "mvn process-resources",
    "configure": null,
    "compile":   "mvn process-classes",
    "test":      "mvn test"
  },
  "resolution": {
    "solved_by": "llm",
    "baseline_result": "fail_at_compile",
    "attempts": [
      { "stage": "A", "toolchain": {"java": "openjdk-21-jdk"},
        "failed_phase": "compile",
        "error_signature": "warnings found and -Werror specified" },
      { "stage": "B", "fix": {"field": "toolchain.java", "to": "openjdk-17-jdk"},
        "rationale": "JDK 21 errors on deprecated-for-removal API; targets JDK 17",
        "result": "success" }
    ]
  },
  "ground_truth_match": {
    "build_system": true,
    "apt_dependencies": { "precision": 1.0, "recall": 1.0 },
    "toolchain_java": true
  }
}

The hypothesis this maps

Porting a technique from one ecosystem to another is engineering, not research, unless there is something about the other ecosystems that makes the problem intrinsically different. I think there is, and it is the hypothesis the study is built around:

Cold-start inference degrades predictably with how much dependency information an ecosystem declares. Maven and Gradle declare their dependencies, so a model can read them. Autotools and Make declare almost nothing, so dependencies only surface as configure or compile errors. Inference should be structurally easier where dependencies are stated and harder where they are latent, and the contribution is mapping that gradient and explaining why.

That also predicts which fields are easy. A build system is declared by the presence of pom.xml or configure.ac, so it should be classified nearly perfectly. A required Java Development Kit (JDK) version is often declared nowhere and only surfaces as a compile error, so it should be recovered poorly by anything that does not actually attempt the build and read the error.

The evidence that the non-scriptable part of this process is large, drawn from the lab dataset:

None of these corrections are recoverable from filenames alone, and none would survive being collapsed into a single Dockerfile.

Prior work, and what I am measured against

Repo2Run builds executable Docker environments for Python repositories, reporting an 86% success rate on 420 repositories [Repo2Run]. CXXCrafter targets C and C++ projects specifically and reports a 78% build success rate [CXXCrafter]. These are the two systems closest to mine, and the comparison design runs my tool and theirs over each other's program sets (mine over their corpora, theirs over the apt-installable subset) so that a reported success rate can be read against a corpus other than the one its own authors chose.

Lyu et al. is the nearest work overall: an LLM generates a working Dockerfile by iteratively patching until the project builds, then optimises image size [Lyu-Dockerfile]. Related work from the same authors fixes missing dependency declarations in existing build files [Lyu-MDfixer], and Rosa et al. trained a transformer to generate Dockerfiles from natural-language requirement specifications [Rosa]. Three differences keep this project distinct:

GradleFixer repairs Android Gradle build failures through domain-specific tools instead of a general-purpose shell, reporting an 81.4% resolve rate on AndroidBuildBench and outperforming a shell-based agent [GradleFixer]. That is where the tool-wrapping part of this design comes from, and testing whether it holds across more than a single ecosystem is one of the research questions.

Two more shape the design without being direct comparisons. Agentless implements a fixed localization–repair–validation pipeline and stays competitive with more complex agents on SWE-bench Lite [Agentless]; that is the logic the deterministic baseline follows. ESAA keeps an append-only log of validated intents for auditability and deterministic replay [ESAA], which is what inspired the recorder here. EChecker [EChecker] and Breaking-Good [Breaking-Good] detect and explain build problems without repairing the environment; they are adjacent rather than competing, and I no longer lean on them to make the case for this project.

The work that started this is Foreman [Foreman], which models per-phase file and behaviour permissions to harden build systems against pipeline poisoning. Its specifications are currently hand-written and automated specification inference is listed as future work. A successful build trace from this project could help draft those specifications, but that is a stretch goal and out of scope for now.

Research questions

Limitations

Stating the limits before being asked is the point, so here they are up front.

Status

Study design settled; environment and harness in progress; results pending. Splitting what is designed from what is running is deliberate. It is the same honesty that makes the evaluation argument credible in the first place.

Settled now

In progress

Still open

What changed after review

An earlier version of this proposal led with documentation and information loss. That framing did not survive review, and the objections were the right ones:

Documentation → build recipe

Documentation is hard to evaluate; there is no defensible way to score whether a document lets someone understand a build. The output is now a recipe with labelled fields, scored on execution in a fresh container and on per-field agreement with the human record.

Rejected: evaluating comprehension or documentation quality.

Multi-ecosystem coverage → the declaration gradient

"It works on more ecosystems" is porting, not a research contribution. The contribution is the claim that inference difficulty is structurally tied to how much an ecosystem declares, and a map that shows where that holds.

Rejected: breadth of ecosystem coverage as the headline claim.

An unsupported premise → dataset numbers

"Most software lacks appropriate documentation" was asserted without evidence. It is now a narrower, measured claim about the programs in the dataset, with the counts stated above.

Rejected: the broad claim about software in general.

References


More work on the projects page, or back to Selected Work.