Explainable Federated Learning

Evaluating and combining explanations for models trained on data held by separate participants.

The problem

In federated learning, participants train a shared model without pooling their data. Explanation methods show which parts of an input drove a prediction, but it is not well understood how they behave when a model is trained this way, or which of several disagreeing methods to trust.

How it works

  1. Train across participants

    Simulated clients train an image classifier on their own share of the data, under different federated algorithms and degrees of difference between the clients’ data.

  2. Explain and score

    Several attribution methods, such as Grad-CAM, SHAP, and LIME, explain the predictions of local and global models, and explanation-quality metrics score the results.

  3. Combine the explanations

    An optimiser weights and fuses the methods’ attributions to improve the metrics while accounting for their computational cost.

What the study found

According to the preprint, the global federated models generally achieved stronger explanation-quality scores than local models across several metrics, and differences between the participants’ data affected explanations more than the choice of federated algorithm. The study also examines how stable the explanation methods are when several similarly accurate models make different decisions. Fusing the methods produced measurable metric improvements.

In a complementary user study, participants significantly preferred explanations aggregated in the federated setting over a baseline trained on pooled data. Some classes were exceptions, so higher metric scores did not always mean explanations people preferred. The paper is a preprint submitted to Information Fusion and has not yet completed peer review.

Using the software

The repository is the replication artefact for the study: a Python package for federated simulations built on Flower, with explanation methods, metrics, and the fusion optimiser. Its reproducibility guide documents the setup and a smoke run; full runs need a Linux machine with a CUDA GPU. The software is MIT-licensed. The work grew out of my master’s thesis, whose raw experiment data and notebooks are archived on Harvard Dataverse.