TraceGate

Choosing useful crash evidence for AI-assisted program repair.

The problem

A repair model needs useful evidence about a failure, but larger prompts can add noise and cost. TraceGate sits between a failing program and an existing repair system, controlling the evidence and budget available to it.

How it works

  1. Capture the failure

    Collect exception details, selected stack frames, local variables, and execution context into a structured snapshot.

  2. Choose evidence and budget

    Lightweight online policies choose what to disclose, whether to invoke diagnostic helpers, and when to increase the repair budget.

  3. Repair and learn

    The existing backend proposes a patch. Execution feedback updates the policies, while a retry gate suppresses redundant patch attempts.

What the study found

Across four Python benchmark suites, five repair backends, and three Qwen3.5 model sizes, the research paper reports recovery of 20.5% of residual failures left after two baseline repair attempts. TraceGate received additional repair budget, so this is not a comparison at equal total cost.

Success means passing the benchmark tests. Most full-suite settings were run once, and policy learning restarted for each dataset, backend, and model. Transfer to other languages or larger, multi-file tasks remains unestablished. The evaluation design and limitations are in Sections VI–IX of the paper.

Using the software

The repository provides installation and usage instructions for the Python package, command-line interface, and integrations. The software is MIT-licensed and retains the package name llmdebug. The Zenodo record provides the evaluation data archive.