Michael Hu
I am a fourth-year PhD student at the NYU Center for Data Science, advised by Kyunghyun Cho and Tal Linzen. I am supported by the NSF Graduate Research Fellowship.
I study the science of training language models, with a focus on data.
Previously, I completed a BSE at Princeton CS, where I spent two lovely years working with Karthik Narasimhan and Tom Griffiths. I then joined Yobi AI for two years as the first employee. In my spare time, I enjoy cooking, running, and climbing.
I’m on the industry job market!
| Sum 2026 | Gave talks on OP-Mix at Cohere, Percepta, and Datology. |
| May 2026 | New paper: Efficient and Simple Data Mixing All The Time (OP-Mix). |
| Feb 2026 | New paper: Neural Neural Scaling Laws. |
| Sum 2025 | Gave talks on Aioli and pre-pretraining at Harvard ML Foundations, MIT Language & Intelligence, and Genesis Molecular AI. |
| Jul 2025 | Pre-pretraining won an Outstanding Paper Award at ACL 2025! 🏅 |
| May 2025 | New paper: Scaling Laws Are Unreliable for Downstream Tasks. |
| Spr 2025 | Gave talks on pre-pretraining at École Normale Supérieure CoML, FLaNN, Ryco Lab Reading Group, and CDS Seminar. |