I'm Cath Ge-Wang, an undergrad at Oxford studying Mathematics. I do technical AI safety research (AI alignment and AI control)
- Languages: Python, JavaScript, C, R, SQL
- ML Frameworks: Pytorch, scikit-learn
- LLM Ecosystem: Hugging Face Transformers, TRL
- Web: HTML, CSS, JavaScript, Flask
- Tools & Platforms: Git, Docker, AWS, Inspect
My primary research interests lie in AI control, evaluations, and model internals, particularly understanding and mitigating emergent misalignment risks in autonomous AI systems. I focus on empirical questions around goal misgeneralisation, alignment faking, and agentic evaluations, aiming to clarify failure modes in dangerous frontier models. I am also interested in how these technical insights inform AI governance and policy, especially hardware verification, mechanisms for strategic risk, and constraining dangerous capability deployment.
Last updated: Jun 2026


