When Weak Intent Becomes a Requirement

Checking whether language models keep how strongly a stakeholder meant a requirement.

The problem

Language models can help turn stakeholder statements into requirements. They may extract the right capability and still change how strongly it was meant: “It would be useful if the system could export reports” becomes “The system should export reports”. Once a wish becomes a commitment, it can carry through into specifications, contracts, and tests.

How it works

  1. Vary only the strength

    360 capabilities from two requirement datasets are each stated four ways: with must, should, may, and as “It would be useful if …”.

  2. Extract the requirement

    Nine models from five model families extract one requirement per statement, with a modality label and a confidence, and are told to preserve the source’s modality. Five further sampled answers show how much the meaning varies.

  3. Check for strengthening

    Wording checks flag outputs that state the requirement more strongly than the source. The study then asks whether confidence, variation across samples, or a separate preservation check catches those cases.

What the study found

Across all source statements, 26.0% of the extracted requirements were stronger than their source under the strict wording check. For weak-intent statements, strengthening was almost universal: 99.7% of readable answers. Most of these (68.9%) became should, shall, or must; nearly all of the rest kept could or may but dropped the wish framing.

The models rarely signalled doubt while doing so. Of the strengthened outputs, 83.4% carried a stated confidence of at least 0.90, and in 67.5% all five sampled labels agreed. A blind check that compared each output with its source flagged 94.8% of them.

These results come with conditions. The wording checks detect specified patterns, so an output they do not flag is not necessarily faithful. The blind check is measured against the strict wording check, not against independently judged semantic errors. The detection scores largely reflect differences between the four kinds of source statement and do not establish a threshold for accepting an extraction or sending it for review. The figures above come from the replication package; the paper is submitted and not yet published.

Using the software

The repository contains the frozen prompts, the benchmark, the campaign configuration, the analysis code, and the provenance behind every reported number; a test keeps the numbers in its README in line with the generated tables. Reproduction is tiered: checks that run on any operating system without model access; recomputing the reported counts, rates, and detection scores from the raw outputs on Zenodo; or rerunning the campaign with your own model access. The software is MIT-licensed.