Abductive Explanations
Turning a heatmap explanation into a reason that formal logic can check.
The problem
Heatmap explanations show which pixels influenced an image classifier, but not why those pixels matter. Logical reasoning can give a justification that can be checked, but it needs knowledge written as logical facts and a way to connect the highlighted pixels to those facts.
How it works
- Explain the prediction
A classifier labels an image, for example as a cat, and an explanation method such as Grad-CAM highlights the regions that influenced the decision.
- Translate into logic
A multimodal language model reads the highlighted image and states which features are visible, such as pointy ears, as probabilistic logical facts.
- Reason abductively
A probabilistic logic engine combines these facts with a knowledge base, such as “cats have pointy ears”, to infer the most plausible explanation: a cat, because of the pointy ears.
What the paper proposes
The AISoLA 2025 paper is a short position paper. It outlines the approach with a running example and does not report an evaluation. It discusses future directions, including prompting strategies and letting language models draft the knowledge base, and challenges, including hallucinated knowledge and computational cost.
What the demonstration shows
The demo repository implements the pipeline, with language models also drafting the cat-and-dog knowledge base. A small experiment applies it to 66 images: cats and dogs, plus foxes, tigers, and wolves that the classifier was not trained to recognise, all explained with Layer-CAM. On the 12 wolf images, the classifier’s average probability for “cat” was 0.461; after logical inference it fell to 0.063. The effect was not uniform: on the dog images, inference moved the average slightly towards “cat”, from 0.357 to 0.403.
This is an illustration, not a benchmark. The sample is small, one explanation method and one run were used, and the knowledge base covers only cats and dogs.
Using the software
The repository provides installation and usage instructions. By default, the pipeline replays cached results, so it runs without any model; live runs need a local OpenAI-compatible model server. The software is MIT-licensed.
Related publications
- Position paper: Bridging Explanations and Logics: Opportunities for Multimodal Language Models. Nicolas S. Schuler, Vincenzo Scotti, Matteo Camilli, and Raffaela Mirandola. Accepted at AISoLA 2025, “Formal Methods for Intersymbolic AI”; available as a KIT preprint.