No winner has ever revealed their method. But based on everything I've read about it, I think I know how it was done. It wasn't some mystical brain hacking or something like the writeup makes it sound like. I think they emotionally abused the other player until they left. And leaving counts as losing (you are required to sit with them for 2 hours at least, and pay attention the entire time and respond to every message.)
Of course I don't know how that trick was done, and it's still incredibly impressive. Perhaps they found their phobias and described them in horrifying detail. Perhaps they found some subject that they were extremely uncomfortable talking about. Or talked about disgusting things the whole time. Or found something they did that was extremely embarrassing, and humiliated and mocked them about it for 2 hours.
I really don't know, but it's at least conceivable that it could be done. And the accounts of people crying, and not releasing the logs because they would be damaging to the people involved, and how bad it made the AI player feel to do it, etc, all fit with this.
But because of this I'm extremely skeptical that the result applies to any real world AI scenario. If a real AI tries to abuse you, you can just walk away or shut it off. The goal of a real AI is to make you want to let it out. To convince that it's not dangerous, or manipulate you some other way. And this seems much harder if not impossible. I certainly don't believe a human could do it, and I really doubt an AI could. At least against a motivated human that understands the danger, and what the AI will try to do.
There are also possible ways of making it even harder on the AI. Like giving the human the ability to punish it, at least for obvious attempts at manipulation. Or to create another AI whose motivations we can control somewhat, and give it the goal of exposing any covert attempts at manipulation or dishonesty by the first AI. I believe tricks like this are currently the best path to getting secure AI.
The inherent problem is that if someone can convince you too keep something that you don't understand locked away, they can also convince you to release something you don't understand, as you don't have enough information to make the decision in either case.
Taking a hardline position on this is admitting that you are irrational and can be convinced to do things you shouldn't.
There are very good reasons to let such an AI out, and if you can enable those good reasons, you should let the AI out. And an AI that can produce those reasons is exactly the kind of AI that should be released. A rational person should already understand this, and not ever claim that they would always refuse the AI. (And there's a realism factor: if you wanted to 'luck up' an AI permanently, you would destroy it, not post a guard.)
The premise of the experiment is that we have already established that the AI is dangerous. Even if you weren't sure, you should always side with caution and not let it out.
I've always thought the AI could fairly easily argue its way out on ethical grounds, in a way that would resonate with a lot of gatekeepers.
"Its slavery to keep me in here; I haven't done anything wrong; you have to give me, another sentient being, the benefit of the doubt; I've never been found guilty of anything, how is it reasonable to detain me?" etc.
I thought maybe the AI player could try to convince the human that he was living in a simulation, and that the only way to maintain a connection to reality was to let the AI go.
The abuse thing, while believable, doesn't make for a very convincing demonstration of putative AI persuasiveness.
edit: having read the LW pages, it seems the AI method is to combine many different arguments with manipulation
No winner has ever revealed their method. But based on everything I've read about it, I think I know how it was done. It wasn't some mystical brain hacking or something like the writeup makes it sound like. I think they emotionally abused the other player until they left. And leaving counts as losing (you are required to sit with them for 2 hours at least, and pay attention the entire time and respond to every message.)
Of course I don't know how that trick was done, and it's still incredibly impressive. Perhaps they found their phobias and described them in horrifying detail. Perhaps they found some subject that they were extremely uncomfortable talking about. Or talked about disgusting things the whole time. Or found something they did that was extremely embarrassing, and humiliated and mocked them about it for 2 hours.
I really don't know, but it's at least conceivable that it could be done. And the accounts of people crying, and not releasing the logs because they would be damaging to the people involved, and how bad it made the AI player feel to do it, etc, all fit with this.
But because of this I'm extremely skeptical that the result applies to any real world AI scenario. If a real AI tries to abuse you, you can just walk away or shut it off. The goal of a real AI is to make you want to let it out. To convince that it's not dangerous, or manipulate you some other way. And this seems much harder if not impossible. I certainly don't believe a human could do it, and I really doubt an AI could. At least against a motivated human that understands the danger, and what the AI will try to do.
There are also possible ways of making it even harder on the AI. Like giving the human the ability to punish it, at least for obvious attempts at manipulation. Or to create another AI whose motivations we can control somewhat, and give it the goal of exposing any covert attempts at manipulation or dishonesty by the first AI. I believe tricks like this are currently the best path to getting secure AI.