<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Generalization on Saurav Panigrahi</title><link>https://sauravpanigrahi.com/tags/generalization/</link><description>Recent content in Generalization on Saurav Panigrahi</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://sauravpanigrahi.com/tags/generalization/feed.xml" rel="self" type="application/rss+xml"/><item><title>Emergent Misalignment</title><link>https://sauravpanigrahi.com/reading/emergent-misalignment/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://sauravpanigrahi.com/reading/emergent-misalignment/</guid><description>&lt;p&gt;Selected references on emergent misalignment and broad behavioral shifts from narrow training signals.&lt;/p&gt;
&lt;h2 id="core"&gt;Core&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href="https://www.emergent-misalignment.com/"&gt;Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs&lt;/a&gt;&lt;br&gt;
Introduces the central phenomenon: finetuning on a narrow harmful behavior can produce broader misaligned behavior outside the training domain.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href="https://www.lesswrong.com/posts/gLDSqQm8pwNiq7qst/narrow-misalignment-is-hard-emergent-misalignment-is-easy"&gt;Narrow Misalignment is Hard, Emergent Misalignment is Easy&lt;/a&gt;&lt;br&gt;
Useful for thinking about why a broad misalignment direction may be a more stable and efficient solution than a narrow one.&lt;/p&gt;</description></item></channel></rss>