Yixiong Hao
yixiong_hao [at] outlook [dot] com
Hi, I'm Yixiong!
I study CS at Georgia Tech. I work on technical research and strategy to reduce societal scale catastrophic risks from highly general, increasingly superintelligent AI systems.
I'm currently a grantmaker at The Astralis Foundation through the Astra Fellowship, where I will be focusing on macro AI strategy and capacity building in Asia. I'm also a co-founder of Second Look Research, and an advisor to the Georgia Tech AI Safety Initiative.
Previously, I
- Ran the Georgia Tech AI Safety Initiative for 2 years. We have produced 15+ peer reviewed publications and helped 20+ people land FTE roles at key orgs such as METR, MATS, Anthropic, IAPS, and Constellation. During my time there, I worked on activation engineering, interpretability, demo-ed AI jailbreaks at a congressional exhibition, and co-authored an RFI response to America's National AI Action plan. I now advise GTAISI.
- Interned at Gray Swan AI, where I analyzed trends in indirect prompt injections and worked on organizational strategy as a pre-deployment evaluator of frontier LLMs.
- Was a research fellow at UChicago XLab, where I worked with CAIS on model organisms of reward manipulation during RL.
Outside of work, I generally love being active. I used to play competitive golf and table tennis. I also like checking out new cafes and restaurants!
Check out some of my amazing friends, collaborators, and mentors. Parv, Will, Zephy, Jasmine, Jo, Andrew, Andy, Rocio, Tzu, Mantas, Animesh, Kartik, Abdur Raheem, Max.
Research
I'm interested in understanding what and how general AI systems learn & generalize in order to develop scalable methods for alignment, evaluations, and control. My latest projects are to do with agentic misalignment in RL with CAIS, and interpretability of VLA models with the PAIR Lab. Research blogs and unpublished work are under writing.