🤖 AI 资讯

· ·
← 返回列表

An Empirical Study of Counterfactual Self-Explanations in LLMs

arXiv cs.CL2026-09-16 04:00:00大模型,算力芯片,Google,Meta,阿里巴巴,扩散模型,模型安全对齐,端侧AI,招聘HR,论文原文 ↗

arXiv:2609.17119v1 Announce Type: new

Abstract: Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation quality: larger models are substantially more likely to generate counterfactuals that flip their own predictions and target decision-relevant evidence. In contrast, the rationale-guided condition produces edit-minimal counterfactuals that are also more human-aligned. However, it does not consistently improve faithfulness. Overall, counterfactual self-explanations can provide useful behavioral evidence about model decisions, but their reliability depends strongly on model capacity and should be empirically validated rather than assumed.