About Me
Hi! I am a Master’s student at KAIST School of Computing, advised by Prof. Alice Oh. My primary research interest is reasoning in large language models (LLMs), with additional interests in post-training and mechanistic interpretability.
My research begins with understanding how reasoning behaviors arise from the internal representations and learning dynamics of LLMs. In particular, I am interested in how these behaviors vary across different operations, languages, cultures, and values. I use mechanistic interpretability to investigate the internal mechanisms underlying such variation and to identify where and why reasoning failures occur.
Building on this understanding, I aim to gradually move from analysis to intervention: developing model-level methods that improve reasoning through pretraining, post-training, alignment, and test-time control. My long-term goal is to build reliable, interpretable, and controllable LLMs that reason robustly across diverse user contexts.
Education
- Korea Advanced Institue of Science and Technology (KAIST) (Sep 2024 - Present)
- Korea Advanced Institue of Science and Technology (KAIST) (Feb 2019 - Aug 2024)
Publications
- *Seogyeong Jeong, *Kiwoong Park, Seyoung Song, Eunsu Kim, Ken E. Friedl, Jaeho Kim, Alice Oh. 2026. LoCar: Localization-Aware Evaluation of In-Vehicle Assistants through Fine-Grained Sociolinguistic Control. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)(ACL-Industry2026)
- *Seyoung Song, *Seogyeong Jeong, Eunsu Kim, Jiho Jin, Dongkwan Kim, Jay Shin, Alice Oh. 2025. MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language. In Findings of the Association for Computational Linguistics: EMNLP 2025 (Long). Also presented at the 5th Workshop on Multilingual Representation Learning (MRL 2025).
- *Junyeong Park, *Seogyeong Jeong, *Seyoung Song, Yohan Lee, Alice Oh. 2025. LLM-C3MOD: A Human-LLM Collaborative System for Cross-Cultural Hate Speech Moderation. In NAACL 2025 Workshop - C3NLP(Workshop on Cross-Cultural Considerations in NLP) [Slides] [Poster]
Preprint
- Haeun Yu, Seogyeong Jeong, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein. 2025. Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models. arXiv preprint
- Paul Röttger, Giuseppe Attanasio, Felix Friedrich, Janis Goldzycher, Alicia Parrish, Rishabh Bhardwaj, Chiara Di Bonaventura, Roman Eng, Gaia El Khoury Geagea, Sujata Goswami, Jieun Han, Dirk Hovy, Seogyeong Jeong, Paloma Jeretič, Flor Miriam Plaza-del-Arco, Donya Rooein, Patrick Schramowski, Anastassia Shaitarova, Xudong Shen, Richard Willats, Andrea Zugarini, Bertie Vidgen. 2025. MSTS: A Multimodal Safety Test Suite for Vision-Language Models. arXiv preprint
Work Experience
Awards & Honors
- Excellence Award (2nd Place), 3rd AI Hackathon for Network Intelligence(Top 2% among 94 teams).
Team award with Jinseo Lee
Teaching Experience
SEOGYEONGJEONG