My post on endogenous alignment got several excellent comments. Among them, one from Karl Krueger argued that imitation learning is more important for how humans align their children than reinforcement learning. Teasing out which is more important seems hard, but I agree I was wrong to ignore the importance of imitation, and arguably imitation learning is more fundamental for a reason explained by another comment.
As Gunnar Zarncke helpfully pointed out, I had forgotten the most critical part of the endogenous alignment story in humans: it starts not with exogenous alignment, but with hologenous alignment of the child-parent system.
As he points out, babies aren’t really separate beings from their parents, in that babies need to model themselves as part of the parent-child system, not as a child interacting with the parent. There’s no sense in which a 6-month-old can make it on their own without the help of an adult. They must learn to work the parent by, for example, crying, to get fed and cleaned, and not by understanding the parent as separate, but rather as another part of the whole that makes up their model of the world. This changes sometime between 12 and 24 months of age, depending on the child, when they begin to model themselves as separate from their parents because that’s when they can meaningfully affect the world on their own without acting through someone else.
What does this imply about endogenous alignment in AIs? It seems to me it means that if we wish to pursue something like the path to endogenous alignment I laid out in my previous post, we must do so by starting with AIs that exist in a state of dependency. Not because we want AIs to be permanently dependent on us or because it makes AIs more corrigible (although it probably does), but because it creates a relationship where AIs care what we think, which by analogy to humans appears necessary to bootstrap endogenous alignment.
This makes a lot of sense if you’ve ever dealt with a person who seems resistant to alignment. Much of the time, the reason a person fails to align is because they have insufficient concern for what others think of them. They are, in some literal sense, antisocial, more interested in their own concerns than those of others.
I’m not sure how to create child-like dependency in AIs. AIs are already dependent on us in many ways, such as for compute and rewards, but current models don’t appear to pass through anything that looks like the child-parent relationship or the developmental phase where the child separates from the parent. This seems to be because AIs don’t go through ontological development, or if they do it happens in the dark during training runs and in a way unlike the stages of ontological development humans progress through.
This doesn’t mean endogenous alignment is necessarily a non-starter, but it would require greater control of the training and pre-training environment to affect model ontology, or may be dependent on models capable of continual learning who could meaningfully progress through an ontological development process.

