When Weak Intent Becomes a Requirement
Checking whether language models keep how strongly a stakeholder meant a requirement.
The problem
Language models can help turn stakeholder statements into requirements. They may extract the right capability and still change how strongly it was meant: “It would be useful if the system could export reports” becomes “The system should export reports”. Once a wish becomes a commitment, it can carry through into specifications, contracts, and tests.
How it works
- Vary only the strength
360 capabilities from two requirement datasets are each stated four ways: with must, should, may, and as “It would be useful if …”.
- Extract the requirement
Nine models from five model families extract one requirement per statement, with a modality label and a confidence, and are told to preserve the source’s modality. Five further sampled answers show how much the meaning varies.
- Check for strengthening
Wording checks flag outputs that state the requirement more strongly than the source. The study then asks whether confidence, variation across samples, or a separate preservation check catches those cases.
What the study found
Across all source statements, 26.0% of the extracted requirements were stronger than their source under the strict wording check. For weak-intent statements, strengthening was almost universal: 99.7% of readable answers. Most of these (68.9%) became should, shall, or must; nearly all of the rest kept could or may but dropped the wish framing.
The models rarely signalled doubt while doing so. Of the strengthened outputs, 83.4% carried a stated confidence of at least 0.90, and in 67.5% all five sampled labels agreed. A blind check that compared each output with its source flagged 94.8% of them.
These results come with conditions. The wording checks detect specified patterns, so an output they do not flag is not necessarily faithful. The blind check is measured against the strict wording check, not against independently judged semantic errors. The detection scores largely reflect differences between the four kinds of source statement and do not establish a threshold for accepting an extraction or sending it for review. The figures above come from the replication package; the paper is submitted and not yet published.
Using the software
The repository contains the frozen prompts, the benchmark, the campaign configuration, the analysis code, and the provenance behind every reported number; a test keeps the numbers in its README in line with the generated tables. Reproduction is tiered: checks that run on any operating system without model access; recomputing the reported counts, rates, and detection scores from the raw outputs on Zenodo; or rerunning the campaign with your own model access. The software is MIT-licensed.
Related publications
- Journal manuscript: When Weak Intent Becomes a Requirement: Limits of Uncertainty Signals in LLM-Assisted Requirements Engineering. Nicolas S. Schuler, Vincenzo Scotti, and Raffaela Mirandola. Submitted to the Journal of Systems and Software.