
I’m a first-year PhD student at Carnegie Mellon University, advised by Prof. Shinji Watanabe. Prior to CMU, I received B.S. in CSE from Seoul National University, and worked at Samsung Research.
My research experience revoles around speech processing, mainly speech recognition and spoken keyword spotting. During my years at Samsung Research, I worked on deployment and optimization of speech technologies for Samsung Bixby, including large-scale training, model compression, long-form speech recognition, and cross-modal inference-time biasing.
In addition to my research in speech processing, I am generally interested in the interpretability of neural networks. I believe that understanding the inner mechanisms of current award-winning large models will help us gain controllability and design more computationally efficient models. On that note, I am interested in the following research topics.
- Speech foundation models: What can we learn from the speech represenations of each layers of the speech foundation models? What is the best way to utilize the phonetic and synthetic information for downstream tasks?
- Interpretability of large models: How are the emergent abilities such as reasoning presented in neural networks? How can we extract the circuits and make them more explicit?
- Controllability and Efficiency: How can we apply our understandings about the neural models to make them more controllable and efficient?
News!
(08/24/2026) I am joining the WAVLab @ LTI/CMU!
(09/24/2024) I will be presenting my paper on spoken keyword detection at INTERSPEECH 2024!