-
Which Traits Resist Subliminal Learning?
Traits more strongly encoded by preference data are harder to overwrite via subliminal learning of the opposite trait, on Olmo and Llama-Tulu after Community Alignment DPO.
-
Can Intrinsic Rewards Save RL?
I argue that prevailing surprise-based conceptions of intrinsic reward fail to explain human curiosity and explanatory understanding.
-
Voting and the Scale-Free Theory of Intelligence
Arrow’s theorem, Popper, and the strange idea that your neurons might be voting too.
-
AI discovers "X"
AI discovers new science (with some help)