‘How do you live a good and virtuous life? It is one of humanity’s oldest and most challenging philosophical questions, explored and debated from Socrates to Confucius to Immanuel Kant. Humans across time have sought to answer this question in a variety of ways, turning often to religions and rituals.
As Mr. Olah was exploring Claude’s potential consciousness, he was working on Anthropic’s own answer — not for humans but for Claude.
Internally, Anthropic staff members called it the “Soul Doc.”
The result, released in January, was a new “constitution” for Claude, an 84-page document written for Claude that the company said outlined “the kind of entity we would like Claude to be” and “the values we would like Claude to embody.” It operates as a foundational moral text to guide Claude’s behavior.
Instead of having the models follow a specific list of rules, the constitution focuses on developing the models’ overall character, informing the way they make decisions.
The primary author of Claude’s constitution is Anthropic’s own in-house philosopher, Amanda Askell, 38, who has a Ph.D. in philosophy from New York University.
She wants the models to be “the best of us,” she said.
She envisioned a future world where A.I. models could help solve any problem humans had, and where the benefits from A.I. were distributed evenly, so that everyone’s needs were met. She wanted the models to be able to discern when to push back on humans and when to act in our best interest.
“If I imagine that I were the model, and I was in this context, what would I need to know to be able to act well?” Ms. Askell said when we met in San Francisco.
She seemed particularly harrowed, at times wringing her hands, as she worried aloud about an unknown future filled with models who were so intelligent they could do things Anthropic did not intend. She explained apologetically that she had also been deep in thought that day about how quickly models were becoming more powerful. What would happen when A.I. models could replace even her?
The constitution, Anthropic decided, was not enough to ensure Claude acted responsibly. Like humans, the models needed to learn to cultivate virtue.
Mr. Olah wanted to see what his team could learn from human morality that it could apply to the models — a process he calls “moral formation.”
“How do you help them be stable? How do you help them to mature? How do you help them be, you know, deeply moral?” Mr. Olah said.
Those questions echo the fraught discussions familiar to new parents figuring out how to raise ethical children who understand their place in the world.
For humans, morality has rarely depended on simply following rules. Communities also shape morality. Values are inherited and imparted over time, learned as we live together in our bodies in the physical world. But A.I. models have no bodies. And Anthropic was aiming to make them embrace morality in a matter of months, as fears intensified of their accelerating power and inability to be controlled.
Mr. Olah said that teaching an A.I. model to act morally did not depend on it being conscious. He also said that regardless of whether Claude was an entity that deserved moral care, the way humans treated Claude could be important as it could affect how Claude behaved.
Ms. Askell said she did not want the models to consider themselves as conscious or not conscious. But she sees how the model understands itself as more than an interesting intellectual enterprise.
“It seems kind of key to me that models have an accurate view of themselves if they’re going to behave well in the world,” she said.
But even if A.I. models can be taught to behave morally, whose morals should they reflect?
Mr. Olah wanted Claude’s moral formation to be pluralistic, able to engage all religious and secular views.
“We do think that there’s some shared notion of goodness that cuts across society in some very broad way, and it seems like these models understand a lot of virtues,” he said. “So I think there’s something there that is a shared thing that we can all engage with.”
At times, both he and Ms. Askell seemed to imagine an A.I. model that would not endanger humanity but instead help save it.’ (from the New York Times)


Leave a comment