Researchers propose C-VCE, a diffusion-based framework for visual counterfactual explanations in vision models
C-VCE integrates a concept bottleneck layer into a diffusion model to generate minimal, interpretable image edits that flip model predictions while preserving visual fidelity.
1 source · cross-referenced
- C-VCE is a diffusion framework that embeds a classifier via a concept bottleneck layer to produce visual counterfactual explanations.
Visual counterfactual explanations aim to answer: what minimal change to an image would flip a model’s prediction? Existing diffusion-based methods can produce realistic edits but depend on external classifiers that must work reliably on noisy images, which limits robustness and deployability in safety-critical settings such as medicine.
Researchers introduce C-VCE, a diffusion framework that integrates a classifier directly into the generative model using a concept bottleneck layer. This design guides counterfactual edits with human-interpretable features (concepts) rather than pixel-level edits guided by a separate noise-robust classifier.
C-VCE allows users to toggle semantic concepts during sampling, minimally adjusting relevant image regions while preserving the rest of the image and respecting feature correlations. A probabilistic regularizer balances the trade-off between changing the prediction and staying close to the original, and a gradient-based mask confines modifications to the most relevant regions.
On benchmarks such as CelebA, C-VCE matches or improves flip rates while producing counterfactuals that are visually closer to the input and less distorted than baselines that rely on separate noisy-image classifiers. The authors argue this makes C-VCE a practical tool for vision systems requiring concrete “what-if” images without relying on an additional noise-robust classifier.
The work suggests that exposing and controlling an internal concept layer is a promising direction for making powerful generative models more understandable and safer to use.
- Jul 28, 2026 · arXiv cs.AI
Researchers propose SeT-Diff, a diffusion-based foundational model for HPC telemetry and time-series
Trust79 - Jul 28, 2026 · arXiv cs.AI
New multi-agent framework uses quantum-classical loops to improve protein structure prediction
Trust79 - Jul 27, 2026 · arXiv cs.CL
Evaluation design alters measured gap between expert and automatic MeSH terms in systematic review classifiers
Trust79