LLM_Deceive_Unintentionally
[ACL26] Experimental resources for paper titled "LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions"
// repository documentation
Was this content helpful?
(0 ratings)
