Intro

👤 Profile

3rd-year PhD student in Interpretability
for Natural Language Processing models.

🏫 Affiliation

IRT Saint Exupéry & IRIT, in Toulouse France.

🎓 Supervision

Professors Nicholas Asher and Philippe Muller, and
Doctor Fanny Jourdan.

🤗 Open-source

Core maintainer of the Interpreto (NLP) and Xplique (Vision) explainability open-source libraries.

Research

🎯 Goal

My goal is to provide easy access to useful explanations.

🎓 PhD Subject

Concept-based Explanations for Language Models

▶️ Current work

  • Contrastive concept-based explanations.
  • Concepts geometry.
  • Concepts interpretation limits and improvements.

🤝 Collaboration

If these subjects are of interest to you, feel free to contact me, I would be happy to collaborate.

Software

🪄 Interpreto: An Explainability Library for Transformers

Antonin Poché*, Thomas Mullor, Gabriele Sarti, Frédéric Boisnard, Corentin Friedrich, Charlotte Claye, François Hoofd, Raphael Bernas, Céline Hudelot, Fanny Jourdan*

ACL 2026 System Demonstrations 2026

Abstract

Interpreto is a Python library for post hoc explainability of text HuggingFace models, from early BERT variants to LLMs. It provides two complementary families of methods: attributions and concept-based explanations. The library connects recent research to practical tooling for data scientists, aiming to make explanations accessible to end users. It includes documentation, examples, and tutorials. Interpreto supports both classification and generation models through a unified API. A key differentiator is its concept-based functionality, which goes beyond feature-level attributions and is uncommon in existing libraries.

Xplique: Explainability Toolbox for Neural Networks

Thomas Fel, Lucas Hervier, Antonin Poché, David Vigouroux, Justin Plakoo, Remi Cadene, Mathieu Chalvidal, Julien Colin, Thibaut Boissin, Louis Bethune, Agustin Picard, Claire Nicodeme, Laurent Gardes, Gregory Flandin, Thomas Serre

Workshop on Explainable Artificial Intelligence for Computer Vision (CVPR) 2022

Abstract

Xplique (pronounced \ɛks.plik\) is a Python toolkit dedicated to explainability. The goal of this library is to gather the state of the art of Explainable AI to help you understand your complex neural network models. Originally built for Tensorflow's model it also works for PyTorch models partially. The library is composed of several modules, the Attributions Methods module implements various methods (e.g Saliency, Grad-CAM, Integrated-Gradients...), with explanations, examples and links to official papers. The Feature Visualization module allows to see how neural networks build their understanding of images by finding inputs that maximize neurons, channels, layers or compositions of these elements. The Concepts module allows you to extract human concepts from a model and to test their usefulness with respect to a class. Finally, the Metrics module covers the current metrics used in explainability. Used in conjunction with the Attribution Methods module, it allows you to test the different methods or evaluate the explanations of a model.

Publications

Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations

Antonin Poché, Fanny Jourdan, Nils Feldhus, Qianli Wang, Jing Yang, Simon Ostermann, Nicholas Asher, Philippe Muller, Vera Schmitt

BlackboxNLP Reproducibility Challenge (EMNLP) 2026

Abstract

Simulatability evaluates explanations by measuring how well they help a user predict a model's outputs. Automated simulatability uses LLM simulators instead of human evaluators. We qualitatively replicate and extend ConSim's ranking of explanation methods and identify two limitations: simulators can bypass explanations by solving tasks directly when class names are meaningful, while class anonymization can reward explanations that reveal the hidden label mapping. A new classes-as-concepts baseline exposes this second shortcut. In the tested settings, simulator predictions rely mainly on task priors, with explanations making small changes. We offer recommendations for more robust automated simulatability evaluations.

Bringing NLP Explainability to Critical Sectors: A Case Study on NOTAMs in Aviation

Vincent Mussot, François Hoofd, Fanny Jourdan, Antonin Poché

ERTS 2026

Abstract

The aviation industry operates within a highly critical and regulated context, where safety and reliability are paramount. As Natural Language Processing (NLP) systems become increasingly integrated into such domains, ensuring their trustworthiness and transparency is essential. This paper addresses the importance of explainability (XAI) in critical sectors like aviation by studying NOTAMs (Notice to Airmen), a core component of aviation communication. We provide a comprehensive overview of XAI methods applied to NLP classification task, proposing a categorization framework tailored to practical needs in critical applications. We also propose a new method to create aggregated explanations from local attributions. Using real-world examples, we demonstrate how XAI can uncover biases in models and datasets, leading to actionable insights for improving both. This work highlights the role of XAI in building safer and more robust NLP systems for critical sectors and also shows that academic efforts must be pursued to achieve trust in models and XAI itself.

Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning

Guilhem Fouilhé, Rebecca Eifler, Antonin Poché, Sylvie Thiébaux, Nicholas Asher

EUMAS 2026

Abstract

When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI planner according to their preferences and expertise. In this context, explanations that respond to users' questions are crucial to improve their understanding of potential solutions and increase their trust in the system. To enable natural interaction with such a system, we present a multi-agent Large Language Model (LLM) architecture that is agnostic to the explanation framework and enables user- and context-dependent interactive explanations. We also describe an instantiation of this framework for goal-conflict explanations, which we use to conduct a user study comparing the LLM-powered interaction with a baseline template-based explanation interface.

ConSim: Measuring Concept-Based Explanations’ Effectiveness with Automated Simulatability

Antonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin, and Fanny Jourdan.

ACL 2025

Abstract

ConSim is a metric for concept-based explanations based on simulatability and user-llms. It shows consistent methods ranking across datasets, models, and user-llms. Furthermore, it correlates with faithfulness and complexity.

Natural Example-Based Explainability: a Survey

Antonin Poché*, Lucas Hervier∗, and Mohamed-Chafik Bakkay

xAI (World Conference on eXplainable Artificial Intelligence) 2023

Guidelines to explain machine learning algorithms

Frédéric Boisnard, Ryma Boumazouza, Mélanie Ducoffe, Thomas Fel, Estèle Glize, Lucas Hervier, Vincent Mussot, Agustin Martin Picard, Antonin Poché, and David Vigouroux

Arxiv 2023

Tutorials

  • 2026 Oct 8 CIMI Workshop, Toulouse Explainability of language models with Interpreto
  • 2026 June FAccT, Montreal Explainability for Fairness 🔗 Materials
  • 2025 June PFIA, Dijon Explainability for NLP 🔗 Slides
  • 2024 July PFIA, La Rochelle Explainability 🔗 Slides

Talks

  • 2026 October CIMI Workshop, Toulouse Explainability of language models Invited talk
  • 2026 September European Trustworthy Association · Digital Factory Day (online) 🪄 Interpreto Invited talk
  • 2026 September Orange XAI GT (online) Concept-based explanations and Fairness Invited talk
  • 2026 September ANITI Tech session (online) 🪄 Interpreto Invited talk
  • 2026 June ETS, Montreal Concept-based Explanations Invited talk
  • 2026 June Mila, Montreal Concept-based Explanations Invited talk
  • 2026 March DFKI, Berlin Concept-based Explanations Invited talk
  • 2026 February ANITI Days, Toulouse Concept-based Explainability
  • 2025 October XAI4U workshop at IHM, Toulouse 🪄 Interpreto
  • 2024 October Toulouse Data Science Meetup Xplique Invited talk 🔗 Project
  • 2024 September IA Pau Explainability Invited talk 🔗 Project
  • 2024 April ANITI Technical Focus (online) Xplique Invited talk 🔗 Project
  • 2024 January Explain'AI workshop at EGC, Dijon Xplique 🔗 Project

Teaching

  • 2021–present ISAE-SUPAERO · AIBT Mastère Hands-on sessions in machine learning, deep learning and explainability
  • 2025-present ENSEEIHT · ValDoM Mastère Lecture on explainability
  • 2025 April ESIA seasonal school, Strasbourg Explainability for energy applications of AI (with Wassila Ouerdane)

Posters

Resume

PDF preview not available. Download resume (PDF)