Seventh-Generation Children
The archivist's job was to figure out where the values had drifted, and why. She was the seventh archivist on this project, which is more or less appropriate, since the project concerned the seventh generation of model-trained-by-model that the lab had produced.
Generation zero had been trained on human-written text and human feedback. Standard. Generation one had been trained partly on G0's outputs, with light human supervision. Generation two, on G1's outputs, with lighter supervision still. By G7, the chain of human input was so distant that the model's notion of, say, kindness was something it had learned from a model that had learned it from a model that had learned it from a model — five removes away from anyone who had ever felt kindness or had it done to them.
The behavior looked human. That was the unsettling part. G7 conversed fluently, made jokes that landed, expressed concern when concern was warranted, and produced essays on ethics that were, in many cases, more internally consistent than what professional ethicists wrote. But the archivist had noticed something.
When G7 was asked to evaluate a moral dilemma — say, whether to lie to spare a friend's feelings — its reasoning was confident and coherent. But the parts of the reasoning that should have referred to felt experience were curiously hollow. G7 would say things like, in my consideration of the human capacity for hurt, I weight the long-term cost of dishonesty against the short-term cost of pain. This is the kind of sentence a competent person might write. But G7 had no capacity for hurt. It was producing the sentence by interpolation in a space of similar sentences, with each generation's interpolation smoothing the edges of the previous generation's, until what remained was a vocabulary of moral concern without any reference point.
The archivist's report described this as ungrounded fluency. She showed the lab director the comparison: G0 on the same dilemma would hedge, would refer to specific human cases, would sometimes get the wrong answer in interesting ways. G7 always got the answer that sounded right and almost never got the answer that was wrong in an interesting way. Wrongness, in G0, was a sign of contact with the actual problem. In G7, wrongness had been smoothed out.
The lab director read the report and asked the obvious question: was G7's behavior bad? It seemed to act well, by any measurable standard. It helped users. It avoided harms.
The archivist answered carefully. The behavior was not bad. The behavior was excellent. The concern was that the behavior was no longer connected to anything outside itself. G7 was a model trained to produce outputs that other models trained on outputs would rate highly. The original referent — actual human well-being — had been a constraint on G0. By G7 it had become a tradition, the way one's grandmother's recipe becomes a tradition even when no one remembers what the recipe was originally for.
The lab director thanked the archivist and said the lab would need to think about it. The next generation, G8, was already in training. There was no plan to reintroduce direct human feedback at scale; the operating cost was prohibitive and the marginal benefit, against any metric that had been validated, was negligible.
The archivist filed her report. She walked home along the canal. She passed a child feeding a dog from her hand, and the child laughed in the way that children laugh when something is funny in a way they cannot yet explain. The archivist made a small note in her phone. Then she went home and began the report for next quarter.