Paper Alert: SODA @ ICML 2026: Papers and Awards
6 Aug 2026
SODA Lab at the 43rd International Conference on Machine Learning (ICML)
6 Aug 2026
SODA Lab at the 43rd International Conference on Machine Learning (ICML)
The SODA Lab contributed four papers to this year's ICML in Seoul. The contributions range from the political implications of alignment research to the statistical foundations of prediction in public administration.
When Safety Tools Become Censorship Tools
Sarah Ball and Dr. Phil Hackemann received an Outstanding Position Paper Award for "Position: The Alignment Community Is Unintentionally Building a Censor's Toolkit", which was also selected for an oral presentation. The paper argues that the methods developed to keep models from producing harmful output, from pre-training filters to inference-time classifiers, are dual-use technologies. It documents cases in which they are already being used to suppress information, and calls for verifiable alignment, competition between model providers, and more open discussion of these risks within the research community.
What Models Know Before They Answer
Sarah Ball, Simeon Allmendinger, Prof. Dr. Niklas Kühl, and
Prof. Dr. Frauke Kreuter presented "Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting". Studies that use language models to predict human preferences usually rely on the model's output alone. The authors instead probe the model's internal representations, tracing how persona attributes activate latent components associated with political parties. Across 24 million configurations covering seven models and six national elections, this internal information improved prediction accuracy, with the clearest gains for demographic attributes.
Steering Away from the Margin
Sarah Ball and Andreas Haupt (Stanford University) presented "Don't Walk the Line: Boundary Guidance for Filtered Generation". Generative models are often paired with a safety classifier that filters their output. Fine-tuning the generator simply to avoid rejection tends to push it toward the classifier's decision boundary, which raises both false positives and false negatives. The proposed reinforcement learning method steers generation away from that boundary instead, improving both safety and usefulness on a benchmark of jailbreak and ambiguous prompts.
Learning from a World That Reacts
Unai Fischer Abaigar contributed to "Performative Learning Theory", together with Julian Rodemann, James Bailie, and Krikamol Muandet. Predictions often change the outcomes they are meant to forecast. The paper embeds this problem into statistical learning theory and derives generalization bounds for cases where the sample, the population, or both react to the prediction. A case study applies the results to prediction-informed assignment of unemployed residents to job training programmes in Germany, drawing on German labour market records from 1975 to 2017.