
Preference Modeling for LLM Alignment under Heterogeneity
Making LLM post-training well aligned with human goals.

Meet the researchers bringing trust and safety in AI from the lab to the real world, live at MIT AI Conference.

Making LLM post-training well aligned with human goals.

We ask how to better emulate rare, extreme events when high-resolution training data are scarce by using the geometry of nudged ensembles to identify local instabilities and inform existing surrogate models non-intrusively.

Reproducing the ELEPHANT LLM sycophancy benchmark, we find its results partly hold but depend on judge and evaluation settings, and that such benchmarks face a short "reproduction window" before the models they evaluate are retired.

Controlling risk while preserving utility in language model generation.

Individualized AI for trustworthy urban evaluation
The Research Slam is a fast, accessible showcase for academics and researchers tackling consequential real-world problems. Presenters distill their work into a memorable story: what matters, what changed, and where it could lead.
This year’s theme, Trust and safety in AI, focuses on how we evaluate, secure, govern, and responsibly deploy increasingly capable AI systems. The stage connects promising research with the operators, collaborators, and investors who can help take it further.