arxiv:2609.01936
🤝 Open to Collab
Matteo He
hematteo
·
AI & ML interests
Generative AI, Mechanistic Interpretability, AI Safety, Post-Training, Reinforcement Learning
Recent Activity
updated a dataset 3 days ago
FedGrok/fedgrok-checkpoints updated a model 8 days ago
hematteo/readout-recipe-control upvoted a paper 8 days ago
Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens