Skip to content
View sagnikc395's full-sized avatar

Block or report sagnikc395

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sagnikc395/README.md

hi, i'm sagnik

i'm passionate about AI safety and post-training alignment of large language models.
my current work focuses heavily on mechanistic interpretability i.e understanding the internal circuits that drive model behavior.

things which i am researching / building stuff , I write over here :
blog · projects

Pinned Loading

  1. qwen2.5-finetuning qwen2.5-finetuning Public

    e2e LORA fine tuning using Unsloth and Qwen-2.5-0.5B

    Jupyter Notebook

  2. dpo-from-scratch dpo-from-scratch Public

    an reimplementation of DPO from scratch

    Python

  3. refusal-circuit-dpo refusal-circuit-dpo Public

    in a tiny-parameter base model, which layers , attention heads, and residual stream directions are responsible for the emergent refusal behavior introduced by DPO

    Python

  4. refusal-steering-using-saes refusal-steering-using-saes Public

    Do SAE Features Outperform Simple Baselines for Detecting and Steering Refusal? [NEMI Workshop 2026]

    Python

  5. rlvr-character-count rlvr-character-count Public

    train a very small language model to accurately count characters in a word in order for it to learn knowledge from verifiable rewards.

    Python

  6. kai kai Public

    building an open source coding assistant that gives amazing context window

    Python