Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

AI Safety & Interpretability Lab

non-profit
https://aisilab.github.io/
aisilab
Activity Feed

AI & ML interests

Interpretability-informed control

Recent Activity

giannor  submitted a paper about 15 hours ago
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
giannor  authored a paper 2 days ago
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
EvilScript  updated a dataset 22 days ago
aisilab/moltbook-files
View all activity

Papers

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals

View all Papers

Lukas Galke Poech's profile picture Stine Beltoft's profile picture William Brach's profile picture Federico Torrielli's profile picture Peter Schneider-Kamp's profile picture Gianluca Barmina's profile picture Filippo Tonini's profile picture

aisilab 's datasets 3

aisilab/moltbook-files

Viewer • Updated 22 days ago • 232k • 57

aisilab/MoltSpeech

Viewer • Updated Jun 2 • 518 • 182

aisilab/moltbook-embeddings

Viewer • Updated May 5 • 189k • 56
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs