Entity

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under ordinary use. We introduce PropMe, a propensity-aware framework for memorization evaluation that contrasts prefix-based capability attacks with non-adversarial evaluations. We propose a metric transformation that, applied to existing functions, allows to create propensity metrics. We further introduce SimpleTrace, a li

Paper · arXiv

cs.CL

Authors: Gianluca Barmina, Peter Schneider-Kamp, Lukas Galke Poech
Published: 2026-06-04
Categories: cs.CLcs.AI

Abstract ↗

via arXiv · 2606.06286