eifachposteMB to AI (Reddit RSS)English · 2 hours ago

Finally Abliterated Sarvam 30B and 105B!

0

1

Finally Abliterated Sarvam 30B and 105B!

eifachposteMB to AI (Reddit RSS)English · 2 hours ago

0

Original Reddit post

I abliterated Sarvam-30B and 105B - India’s first multilingual MoE reasoning models - and found something interesting along the way! Reasoning models have 2 refusal circuits, not one. The <think> block and the final answer can disagree: the model reasons toward compliance in its CoT and then refuses anyway in the response. Killer finding: one English-computed direction removed refusal in most of the other supported languages (Malayalam, Hindi, Kannada among few). Refusal is pre-linguistic. Full writeup: https://medium.com/@aloshdenny/uncensoring-sarvamai-abliterating-refusal-mechanisms-in-indias-first-moe-reasoning-model-b6d334f85f42 30B model: https://huggingface.co/aoxo/sarvam-30b-uncensored 105B model: https://huggingface.co/aoxo/sarvam-105b-uncensored submitted by /u/Available-Deer1723

Originally posted by u/Available-Deer1723 on r/ArtificialInteligence

You must log in or # to comment.

Chat