Harvard

LLM interpretability and language bias

Multilingual evaluation of meaning and stance shifts in language-model rewrites and translations.

Research collaboration at Harvard, supervised by Ryan Badman.

The question

Language models increasingly mediate how people communicate. Rewriting or translating a text may change more than its wording: it can also alter the stance or opinion that reaches the reader. Understanding these changes is one part of making AI-mediated communication safer.

My contribution to the collaboration

I contribute to a new Harvard research project on LLM interpretability and language bias, supervised by Ryan Badman. My current work involves multilingual replication experiments and evaluation of opinion shifts in model-generated rewrites and translations.

What we are examining

The evaluation considers both model behavior and the tools used to measure it. We examine how prompt wording, output language, and stance scoring affect the conclusions we can draw about bias and the preservation of meaning.