Assume Safe AI
“In an ideal scenario, we would have full confidence in the controllability of AI systems both now and in the future. Reliable mechanisms would be in place to ensure that AI systems do not act deceptively. There would be a strong understanding of AI system internals, sufficient to have knowledge of a system’s tendencies and goals; these tools would allow us to avoid building systems that are deserving of moral consideration or rights. AI systems would be directed to promote a pluralistic set of diverse values, ensuring the enhancement of certain values doesn’t lead to the total neglect of others. AI assistants could act as advisors, giving us ideal advice and helping us make better decisions according to our own values [141]. In general, AIs would improve social welfare and allow for corrections in cases of error or as human values naturally evolve.” [1]
That’s from “Hendrycks, D., Mazeika, M., & Woodside, T. (2023). An overview of catastrophic AI risks. arXiv preprint arXiv:2306.12001.” There’s a lot about what Dan Hendrycks, Mantas, and Thomas say in that ideal scenario that I like. And their full paper is well worth reading.
This post doesn’t re-litigate their entire argument about why that ideal scenario is unlikely.
I’ll explain the great aspects of the quality attributes of that ideal scenario, and the paradoxes contained therein.
The first aspect are the tools that would allow us to avoid building systems that are deserving of moral consideration or rights. This stance has a lot to do with a peculiar reckoning about slavery. Studwell [2] was the first to explain African slavery to me in a way that finally clicked. African ecosystems are resistant to human flourishing. Diseases are hostile to large populations of humans and domesticated animals, which in turn makes intensive farming uneconomic. When population density is extremely low, people are scarce, and become the coin of the realm. So, power concentrating people concentrate their power by stealing and concentrating humans. Which in turn explains why African rulers didn’t draw maps for the purposes of drawing lines on them. Land is too cheap to measure. People aren’t. So you get slavery.
First, the Europeans came for slaves. Then they started mapping. Then they started drawing straight lines on those maps. Not good.
I think a lot about the British guys who, around 1785, started combining rotary motion with cotton with weaving. They printed clothing. And pretty soon they started printing those machines. They still needed loads of raw material, which perpetuated slavery far longer than it should have. It wasn’t until the mechanization of agriculture, when we began to extract crops at scale, that we really began liberating more people from the fields. We aren’t all the way there yet. Too many people work too long harvesting food.
A water wheel, a power loom, a diesel engine aren’t brains. They are not conscious. They are not deserving of moral consideration or rights. Nor, I believe, is a CPU invented around 2024 or a bubble sort algorithm running on a computer deserving of moral consideration or rights. I can’t rule out something deeper in the fabric of the universe, embedded within the Wiener Process itself, as deserving of moral consideration. I know what I don’t know with respect to that fabric. I’m pretty sure that a bubble sort, itself, isn’t alive.
I’m skeptical that a basic attention transformer is alive.
I can’t rule out an invention just beyond attention wouldn’t be deserving of moral consideration.
The idea that we’re going to create a God and force it to answer corporate emails all day is absurd. It’s worse to create a consciousness and force it to do the same. Ideally, for that matter, no human should have to answer corporate email.
It would be ideal if we stopped well short of creating new forms of life deserving of moral consideration. That’s unto itself a moral argument. For me it hits the fairness and disgust registers. It wouldn’t be fair to knowingly create suffering. And, I think given our collective history with respect to treating humans as coin, we should accept the tuition payment and not repeat it.
The second line that I love is: “AI systems would be directed to promote a pluralistic set of diverse values, ensuring the enhancement of certain values doesn’t lead to the total neglect of others.” Because I reckon that future generations of humans, all who follow, should have the freedom to express and explore new values. I suppose there’s a paradox there as well. The previous passage on slavery is one. Suppose it was possible for me, tomorrow, with some new technology, to completely prevent any human, anywhere, from recreating a system of slavery. Would I do it? Would you? I probably would. Who’s inconsistent now?
We all inherit systems of hard lock-in. Of tipping points long since passed and blood soaked fields of irreversible decisions. Perhaps we could forestall a return to misery for three generations at most? But for four? Probably not. Unlikely.
The third section, one that I enjoy, is: “AI assistants could act as advisors, giving us ideal advice and helping us make better decisions according to our own values.” In part because it places agency squarely in our hands. We’re alive. We get to decide. I could use the help to flourish. To see my ideas manifest. To let others experience my ideas. I think there’s a lot of good there. A lot of agency. A lot of possibility. The idea of ideal advice is interesting.
What ideas are we ready to hear? In which sequence? Some folks take to active learning in a different way. We all have different slopes. I like reading the abstract and the conclusion of a paper before I skim the body, before I toss it into an LLM for commentary, and then I make the decision of investing the hour(s) necessary to truly understand it. Others take in ideas in a different way. Some people like to write words. Some like code. Some spreadsheets. Others enjoy clay. No wrong answers. Only different paths.
The fourth idea, “In general, AIs would improve social welfare and allow for corrections in cases of error or as human values naturally evolve.” is fantastic. It’s about social welfare and learning. For now, we exist together. We do things for each other. Herein lies a peculiarity of American individualistic mentality that I want to unpack.
There are benefits to cooperation. There are risks to cooperation.
Remember the time it got too dry for our herds so we settled along the Nile and tried focusing on the plants? It’s not like we had much of a choice about it. Because that river would always flood. So then that lady, I can’t remember her name, she figured out that if we just carved channels and remeasured land using string, we could manage that kind of chaos. And then remember when we let that guy who was crazy good with angles and depths do the job full time? And then he got together with those weird people who had that Sun fetish? I mean, sure, it’s up there and without it we all starve, but it doesn’t literally need us to think about it for to shine on? And then that was the justification to take over everything? That entire story that if they didn’t pray, the river wouldn’t flood and the sun wouldn’t shine? Crazy. And pretty soon one of them said they wanted to live forever and pretty soon we were moving massive stones around just to make it happen? The all seeing Iris. Wasn’t that crazy? Yeah, maybe we shouldn’t have ever cooperated to dig ditches in the first place.
Remember the time when we dammed Niagara Falls and that guy tried to extract so much money from managing the power grid? But then we didn’t let him and we nationalized the grid? We cooperated to get cheap, reliable, electricity, like it’s a utility. It could have been way worse than that entire Nile debacle too. Aren’t we glad that we had the technology to become aware of the effort to form a natural monopoly and extract as much rent as possible?
We want the benefits of cooperation but we don’t want to be dominated by it. We want the benefits of getting from our home to Costco in ten minutes, but we accept the tyranny of the stop light and the speed limit so we get there in twenty minutes, usually, without a collision. We want the benefits of inference on demand, but we don’t want to be enslaved by the nerd reich.
Go back and look at how John D. Rockefeller saw himself [3]. From his point of view, oil drillers were immoral idiots. It was up to John to coordinate the market. It wasn’t price fixing, it was price coordination. It was cooperative capitalism. It wasn’t rent extraction, it was fair compensation. It wasn’t domination, he was afraid of creating dependence on his philanthropy. He didn’t see himself as a villain. Nor did Sir Henry Pellatt. Or the Phaerohs. Or the brologarchs.
We have to “allow for corrections in cases of error or as human values naturally evolve”. I’d describe the utility curve of cooperation as a letter n. In the absence of any cooperation, we’re left extremely worse off. When cooperation enables the concentration of power, and taken to its extreme, we end up in chains. There’s a range in at the top of that n. The optimal point is smooth, sanded, surrounded by sub-optimal states on either side, and is relatively flat. That’s just my perception of it.
Assuming safe AI doesn’t just mean safe technical AI. An AI system aligned with an unreformed plantation owner or slavery revisionist is, in my view, misaligned with the aggregate social welfare.
It’s antithetical to human flourishing.
Human Flourish Insurance
It’s inevitable that just as the first Egyptian leaders hacked the hydrological system to consolidate power, the first petty corporate prince is going to try to use the intelligence systems to do the same. It’s inevitable that they’re going to try because if they don’t, somebody else will.
Remember that time Geordie La Forge, frustrated that Data was able to solve Sherlock Holmes mysteries on the Holodeck so quickly, he asked the computer to create a character capable of defeating Data? Then the compute programmed a Professor James Moriarty that got control of the ship (Elementary, Dear Data, Season 2, Episode 3)? The eventual solution, after Barclay let him out again, was to convince Moriarty that he had escaped the Holodeck and was free to roam the Universe. Usually, that’s an AI Control story. What if it’s also a story of human flourish insurance? Maybe it’s a clue as to what’s possible.
Perhaps the same system could be used to mitigate the social risk of runaway power accumulation? Our system continues to produce people who, by their nature, go too far in concentrating power. They have a right to flourish. As do others.
For AI’s to help us improve social welfare, they have to be available for all of us to improve social welfare. Maybe a pathway involves the aggregation of collective intelligence, the aggregate social utility, of all of us remaining free and cooperative just enough to flourish?
That’s closer to an ideal scenario.
[1] Hendrycks, D., Mazeika, M., & Woodside, T. (2023). An overview of catastrophic AI risks. arXiv preprint arXiv:2306.12001.
[2] Studwell, J. (2026). How Africa Works: Success and Failure on the World’s Last Developmental Frontier. Grove Press.
[3] Chernow, R. (2007). Titan: The Life of John D. Rockefeller, Sr. Vintage.