Skip to content
View GusSand's full-sized avatar

Highlights

  • Pro

Block or report GusSand

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
GusSand/README.md

Gustavo Sandoval

I work on the gap between how safe a model looks and how safe it is — from the inside and under attack. Mechanistic interpretability and LLM security. Ph.D., NYU Tandon (2026); before that, 15+ years shipping software at Microsoft.

Selected work

Elsewhere

gussand.github.io · Google Scholar · @gussand

Pinned Loading

  1. marin_red_teaming marin_red_teaming Public

    Independent Red Teaming for the Marin models. (https://github.com/marin-community/marin)

    Python

  2. surgeon surgeon Public

    Code for "Locating Is Not Repairing": the same decimal-comparison bug arises from a different attention-head circuit in each LLM family, and published fixes fail to transfer

    Python

  3. CoT_Exploration CoT_Exploration Public

    Mechanistic experiments on latent chain-of-thought (CODI): what happens to reasoning, and to its interpretability, when explicit CoT is compressed into continuous space

    HTML 1

  4. MATS MATS Public

    Code and experiments for "Even Heads Fix Odd Errors" (arXiv:2508.19414): attention-head analysis and surgical repair of a format-dependent numerical comparison bug

    Python

  5. research-radar research-radar Public

    Daily + weekly automated radar of new mech-interp & AI-security research (arXiv, OpenReview, ACL, TMLR, LessWrong/AF). Written by Claude Code cloud routines.

    HTML