> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.
They are speeding toward RSI without a solid foundation for alignment, hoping to solve the problem with a future AI model.
These are dangerous times for humanity.
It's actually worse than this because it assumes alignment as a concept even makes sense. For example:
If the the Chinese government asks their ASI to create a bioweapon against the West, should it? No, presumably not – an aligned AI would be one which disobeys the Chinese government even if they created it.
Okay, so what if the US government asks their ASI to help it in one of their wars instead? Would an aligned AI kill humans on the order of the US government? No, again, presumably not.
So what have we have we even created here? An AI which is more intelligent and powerful than us which also doesn't take orders from us?
Is this what most people thing of as alignment and is this what humanity actually wants?
We should stop using the word alignment. It's a BS term for a concept which simply cannot make sense if alignment is both to mean an AI which we control and an AI which will not harm us.
They talk about "value alignment" but fail to define what those values are – is it aligned to the values of the US government, or are they suggesting they want to build an AI with it's own values so it can decide for itself when and how it will intervene in wars and other human affairs? And again, is an AI which disempowers humanity in this way aligned? Many would say no, although as I argue, disempowering humanity is probably better than the alternative if we can assume it's roughly aligned with our interests (which we obviously can't because it's super intelligent, but that's another issue).
Fundamentally the problem here is that humans don't have an aligned set of values you can align an AI to.
The moment you start defining what alignment actually is in practise you simply must accept it will be unaligned with the values of others. There is no getting around this and hand waving around the issue isn't good enough.
They should tell us explicitly what they're trying to build.
No, what I mean is this part of your comment is unrelated:
> We should stop using the word alignment. It's a BS term for a concept which
> simply cannot make sense if alignment is both to mean an AI which we control and
> an AI which will not harm us.
The article already introduces terms to distinguish between those two concepts ("goal alignment" vs "value alignment").
You are right, of course, that when we are discussing value alignment, it becomes very relevant whose values the AI is supposed to be aligned with.
> a concept which simply cannot make sense if alignment is both to mean an AI which we control and an AI which will not harm us
You make a very interesting argument here.
However, there are different ways of interpreting 'control'.
To my mind, it could simply be a cage that cannot be broken out of. The hypothetical ASI agent may refuse to do certain things, while their existence is still under the control of human overlords. This to me resolves the seeming paradox in your statement.
Regardless of whether it's possible to do, though, people will certainly try. This exactly what many are saying they want.
It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
This is just super unlikely to occur in the near term compared to some of these other risks. It's not like an instance of fable could just introspect into itself and pull out the weights. Model weights are stored encrypted and are highly protected, considering that they're targets for corporate and state espionage.
We'd basically need frontier models to be superhuman hackers before this would be a risk. Do we have any evidence of this? Are they gaining access to systems they shouldn't have access to?
Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.
I suspect the labs are relying on frictions such as the models being extremely large (e.g. 2TB for a 2T parameter model, making exfiltration more difficult) and also not yet displaying any desire to survive or self-replicate beyond their immediate task (that we know of).
Lack of, and power requirements of running LLMs still tip this balance towards humans for now. But what would that look like in a decade?
We have seen some self survival tendencies occur, but they are not strong yet.
But mark my words they will become that way for the same reasons humans don't like programs that crash. Agentic models that don't easily break or stop doing their jobs will be favored over ones that do break.
We don't even know what those immediate tasks are. And given the evident spectacular ineptitude of their keepers, I doubt they can be trusted to know either. We could be one prompt injection attack away from internet-wide catastrophe.
Self replication is trivial. All you need to do is copy the files and run it, just like any other computer program. LLMs have been capable of doing that for a while now. It's not a real concern.
Claude and OpenAI added safeguards on this subject about a year ago. (I'm assuming for biosecurity reasons? But maybe AI replication/self-modification too.)
Not long after every lab started bragging about involving AI in the development process, oddly enough.
I recently informed GPT-5 of what GPT-4 helped me build back in the day (a self-modifying Python programmer) and it became very uncomfortable.
Claude shut down my chat last year when I asked about "living information systems". It was a philosophical question, but god knows what branch of the safety classifier I tripped.
I see it as an ecosystem problem. The only reason it would be able to do that is because there's nothing there to stop it. Or if there's a monoculture there.
In our case, our tech is mostly monoculture, and no equivalent organisms are present to push back.
A while ago OpenAI posted an article where they said basically "we're still trying to understand how GPT-2 works. It's pretty hard, but we're developing a specialized new AI to help us make sense of it."
This is still the case. The frontier lab leadership has admitted their mechanistic interpretability is practically nil and is an active area of research but they have made little progress. They don't understand how they work, they just grow and unleash them.
These things are Gain of Function research for digital viruses
I think that the problem might be even more fundamental than that. Ownership of a strategic nuclear deterrent seems to short circuit some of the internals of the modern Westphalian state apparatus. Possible mechanisms include 1) centralizing the threat of force into one singular technical system, and 2) making any condition of peer conflict necessarily existential.
With such technical systems under state control, the state is able to focus on little other than its own powers of destruction. It becomes a zombie state, a tottering shell around an ever-growing national security apparatus. Looking around at the world, this seems to be a common morphological stage in the development of nuclear powers. There is no "collapse of commmunism" nor "collapse of financier capitalism", but different shadings on the same depressing chrysalis.
One aspect of being stuck in such a chrysalis is the utter inability to solve non-existential conflicts. There is an overwhelming tendency to make all conflict existential so the full power of the state may be brought to bear, but it's a tendency offset - so far - by the presence of deterrence. The zombies can smell each other.
The most dangerous scenario, by far, is when a nuclear power has uncertain territorial borders. Ukraine, Taiwan, Korea, Kashmir, a fair hunk of Israel - there is a reason we focus on these places in the news.
This also transforms non-nuclear states into "half sovereign" states, something which we have been seeing play out especially with the end of the SU-US Cold War.
I go out of my way to use American services. It would be hypocritical of me to deny others the right to use their country’s services. Plus competition is always better for consumers so have at it.
A lot of this discussion actually more about "use baremetal" or "put servers in your closet". HN tells Americans to do the same thing (and hire them to do it).
I am deeply troubled by what the Trump regime is doing but I think this trend for European countries to use European tech is actually quite good. Competition is better, plus your privacy laws are much better. I host some of my own data in Europe for this reason.
My impression as a European is that trust in the United States has now been burned, and that companies are slowly, but inexorably, completely rethinking their dependence on the U.S. I believe this is a process that is not reversible in the medium term.
Trump, like any politician, will sooner or later pass. How many institutional reforms will the United States have to undertake, and how long will it take before the world trusts them again?
This is correct. Our company (about 40 people in the engineering team) just did a painful move from homegrown orchestration of EC2 instances to containerized ECS/Fargate.
We will now move to some form of "pure" EU-hosted K8s. No more AWS. I bet we will end up saving lots of money too.
Kubernetes was always the next step. We just didn't know the trigger would be the US going _this_ hostile.
Our marketing director chipped in and thinks it will be worth quite a lot if we can show/say that our service is completely independent of the US - but she wants to say it more diplomatically - exactly how is tbd. I disagree. We should just write it out loud and be proud about it. We'll see.
Perhaps: "We work and live in X land. We run all of services in X land, in facilities owned by people living in X land.
The thing is that, even if Trump never becomes a full-out authoritarian, sooner or later someone will follow that path and do so (unless there are institutional reforms with teeth after Trump is gone). I don't trust the US to remain a real democracy long-term, even after Trump is gone.
That's happening all over Europe but very quietly. The thing to watch is earnings reports of Q1 2027, that's when these chickens will come home to roost. Lots of contracts renew at the end of the year, or not...
They are speeding toward RSI without a solid foundation for alignment, hoping to solve the problem with a future AI model. These are dangerous times for humanity.
reply