Yixiong Hao郝奕雄
yixiong_hao [at] outlook [dot] com
Hi, I'm Yixiong!
你好,我是奕雄!
I study CS at Georgia Tech. I work on technical research and strategy to reduce societal scale catastrophic risks from highly general, increasingly superintelligent AI systems.
我在佐治亚理工学院(Georgia Tech)学习计算机科学。我从事技术研究与战略工作,致力于降低高度通用、日益趋近超级智能的 AI 系统所带来的社会规模灾难性风险。
I'm currently a grantmaker at The Astralis Foundation through the Astra Fellowship, where I will be focusing on macro AI strategy and capacity building in Asia. I'm also a co-founder of Second Look Research, and an advisor to the Georgia Tech AI Safety Initiative.
我目前通过 Astra Fellowship 在 Astralis Foundation 担任资助官(grantmaker),专注于宏观 AI 战略以及亚洲地区的能力建设。我也是 Second Look Research 的联合创始人,并担任佐治亚理工 AI 安全倡议(Georgia Tech AI Safety Initiative)的顾问。
Previously, I
此前,我
- Ran the Georgia Tech AI Safety Initiative for 2 years. We have produced 15+ peer reviewed publications and helped 20+ people land FTE roles at key orgs such as METR, MATS, Anthropic, IAPS, and Constellation. During my time there, I worked on activation engineering, interpretability, demo-ed AI jailbreaks at a congressional exhibition, and co-authored an RFI response to America's National AI Action plan. I now advise GTAISI.
- 领导佐治亚理工 AI 安全倡议(GTAISI)两年。我们产出了 15 篇以上同行评审论文,并帮助 20 多人在 METR、MATS、Anthropic、IAPS、Constellation 等重要机构获得全职职位。在此期间,我研究了激活工程与可解释性,在一次国会展览上演示了 AI 越狱,并合著了一份针对美国国家 AI 行动计划的 RFI 回应。我目前担任 GTAISI 的顾问。
- Interned at Gray Swan AI, where I analyzed trends in indirect prompt injections and worked on organizational strategy as a pre-deployment evaluator of frontier LLMs.
- 在 Gray Swan AI 实习,分析间接提示注入(indirect prompt injection)的趋势,并作为前沿大模型的部署前评估方参与组织战略工作。
- Was a research fellow at UChicago XLab, where I worked with CAIS on model organisms of reward manipulation during RL.
- 在芝加哥大学 XLab 担任研究员,与 CAIS 合作研究强化学习中奖励操纵的模型生物(model organisms)。
Outside of work, I generally love being active. I used to play competitive golf and table tennis. I also like checking out new cafes and restaurants!
工作之外,我喜欢运动。我曾经打过竞技高尔夫和乒乓球。我也喜欢探店,尝试新的咖啡馆和餐厅!
Check out some of my amazing friends, collaborators, and mentors. Parv, Will, Zephy, Jasmine, Jo, Andrew, Andy, Rocio, Tzu, Mantas, Animesh, Kartik, Abdur Raheem, Max.
来认识一下我那些很棒的朋友、合作者和导师:Parv、Will、Zephy、Jasmine、Jo、Andrew、Andy、Rocio、Tzu、Mantas、Animesh、Kartik、Abdur Raheem、Max。
Research研究
I'm interested in understanding what and how general AI systems learn & generalize in order to develop scalable methods for alignment, evaluations, and control. My latest projects are to do with agentic misalignment in RL with CAIS, and interpretability of VLA models with the PAIR Lab. Some of my academic services include serving on the AAAI alignment track senior program committee, the Longitudinal Expert AI Panel (LEAP), and reviewing for workshops at ICLR, ICML, NeurIPS, and EMNLP. Research blogs and unpublished work are under writing.
我希望理解通用 AI 系统学到了什么、如何学习并泛化,从而开发可扩展的对齐、评估与控制方法。我最近的项目包括与 CAIS 合作研究强化学习中的智能体失调(agentic misalignment),以及与 PAIR Lab 合作研究 VLA 模型的可解释性。我的学术服务包括担任 AAAI 对齐方向的高级程序委员会委员、纵向 AI 专家小组(LEAP)成员,以及为 ICLR、ICML、NeurIPS 和 EMNLP 的研讨会审稿。研究博客和未发表的工作见写作页面。