AIs have passed the Turing Test

That’s a possibility, I must admit. But why would the internet be flooded with the obvious kind of AI, if the “really good stuff” is available?

My guess is that the “really good stuff” will require a new, different kind of technology than just larger and larger large language models.

I think we’re already at the point of diminishing returns with this particular technology. They’ve already exhausted (and/or poisoned) the internet’s collected information, and now are destroying old books at a breakneck pace to squeeze any improvement out of these algorithms.

Better yet, countries with nukes, including the U.S., are working on integrating AI into their nuclear weapons systems. But don’t worry, it’s not as if we’re going to give HAL the responsibility to launch a nuke on its own, so what could go wrong?

https://thehill.com/opinion/national-security/6058469-autonomous-weapons-global-security-threat/

Every time I think, “People can’t be that stupid,” I am proven wrong.

The whole business with AI has just been the nail in the coffin of my humanitarian idealism.

This is correct. It was not a definition of consciousness or of unimaginable capability. It was just a thought experiment about how one measure the vague concept of “intelligence”. Not superintelligence or consciousness or anything like that.

It was just like, if we ever build computers much more powerful than we have in 1950, we’ll need to assess intelligence, and here’s one basic bar we can set.

AFAIK, software has passed Turing tests since long before AI, but the test kind of falls apart once participants have some exposure to the system’s behaviors and limitations. 3 years ago, an LLM could’ve fooled me on a Turing test. Now that I have enough familiarity with them, they wouldn’t be able to fool me.

That’s not to dismiss the sophistication of AI but rather to highlight the limitations of the Turing test, as well as the complexity of the questions of intelligence or consciousness. I think in many ways AI’s have more smarts than many people, and in situations on where your interactions with the AI are constrained to certain statements or questions, then it could perform as well or better than a person, maybe. But in open-ended scenarios, I don’t see that happening anytime soon.

I learned of the Turing Test long before there were any AIs who could come close to passing it, and this is exactly my conception of it.
That LLMs can kind of pass it today might mean that LLMs are smarter than we think or that people are dumber.
I know enough about cognitive biases to fool someone into making an irrational decision, but that doesn’t mean the person is not intelligent, whatever that means. Neither does manipulating AIs through a knowledge of how they work.

It’s really just a limitation of the test. Different kinds of software have been passing the Turing test (for definitions of “passing” at the time) for decades now. We blew through the 1950’s bar decades ago, and we’re now just learning all the different gradations of “passing”.

For example right now if I were sitting on a blind Turing test panel, I’d have to challenge it in different ways than I would 20 years ago. I’d ask it detailed technical questions in 2 entirely unrelated fields, and if it answered both of them fluently rather than saying “that’s not my area of expertise” on one of them, then I’d know it’s AI. Or ask it questions about its life and backstory including a lot of chronlogical sequencing, which is something I’ve observed AI have a really hard time keeping their story straight. There are all sorts of tells that could be used.

Putting out a lot of technically detailed content is something that would’ve passed for intelligence maybe 15 years ago, but now we know that the lack of ability to curate or constrain what’s returned is a hallmark of artificial intelligence. The Turing test, again, was never intended to test intelligence at this level, but rather to test what was considered well out of reach in 1950

I used “Turing Test” as shorthand for the headline, to indicate that AIs now of capable of many human-like behaviors that few would have believed even a few years ago. The responses here seem to fall into a group saying that no machine can ever pass a Turing Test and another saying that passing a Turing Test is easy but doesn’t really measure anything.

It might be helpful to go back and look at Turing original paper from 1950, A. M. Turing (1950) Computing Machinery and Intelligence. Mind 49: 433-460.

He starts with “I propose to consider the question, ‘Can machines think?’” This was not a new or startling question. People were already talking about it, in ways startlingly similar to what I’m seeing here. Most of their responses to the question was “no.” Turing’s paper gave a stripped-down way of possibly finding a way to answer “yes,” the Imitation Game.

By modern standards, the examples are trivial for an LLM.

Q: Please write me a sonnet on the subject of the Forth Bridge.
A : Count me out on this one. I never could write poetry.
Q: Add 34957 to 70764.
A: (Pause about 30 seconds and then give as answer) 105621.
Q: Do you play chess?
A: Yes.
Q: I have K at my K1, and no other pieces. You have only K at K6 and R at R1. It is your move. What do you play?
A: (After a pause of 15 seconds) R-R8 mate.

Note that most humans would fail these questions, especially if they had to do so in their heads. No machine of that time could write a sonnet then, either, but they can and do today. I’ve never written a sonnet in my life, even though I’ve written millions of words.

A sonnet requires formal knowledge of composition. A “no” arguer would require more. Turing quotes Professor Jefferson’s Lister Oration for 1949,

“Not until a machine can write a sonnet or compose a concerto because of thoughts and emotions felt, and not by the chance fall of symbols, could we agree that machine equals brain-that is, not only write it but know that it had written it. No mechanism could feel (and not merely artificially signal, an easy contrivance) pleasure at its successes, grief when its valves fuse, be warmed by flattery, be made miserable by its mistakes, be charmed by sex, be angry or depressed when it cannot get what it wants.”

If fed into an AI (and the word valves replaced), it would not be to tell if that were written in 1949 or 2026. No human could either.

Turing admits that it would “idiotic” to insist that a machine could enjoy strawberries and cream, another contrary example, but basically says this is a distinction without a difference. Humans have a range, machines will have a range. If perfect overlap is required then the answer must be “no,” machines are not humans. But that doesn’t really address the question of “do machines think?” Machines do not need to be all-knowing, perfect, or even sound like humans to be capable of thinking.

I would add that no one would ask that of all humans. If machines must not merely precisely match humans but exceed them in every way to be capable of thinking there is no question to be argued. That precludes the possibility of a different way of thinking a priori. I don’t think Turing would accept that.

He also short-circuits an attitude much seen today, here and elsewhere.

“The consequences of machines thinking would be too dreadful. Let us hope and believe that they cannot do so.”

This argument is seldom expressed quite so openly as in the form above. But it affects most of us who think about it at all. We like to believe that Man is in some subtle way superior to the rest of creation. It is best if he can be shown to be necessarily superior, for then there is no danger of him losing his commanding position. The popularity of the theological argument is clearly connected with this feeling. It is likely to be quite strong in intellectual people, since they value the power of thinking more highly than others, and are more inclined to base their belief in the superiority of Man on this power.

I do not think that this argument is sufficiently substantial to require refutation. Consolation would be more appropriate: perhaps this should be sought in the transmigration of souls.

He ends by saying, “We can only see a short distance ahead, but we can see plenty there that needs to be done.”

We’ve done plenty. Machines are not humans, but it’s becoming increasingly ridiculous to acknowledge the possibility of intelligent aliens but not intelligent machines.

In some far-flung future, maybe, but LLMs are not intelligent. This must be evident to anyone who uses one. My experience of Claude is that it doesn’t understand anything and seems to have no coherent framework of reality. It gives an answer and then provides citations that have nothing to do with its answer. It doesn’t know what day or time it is. I use it for managing my grants workflow and it frequently conflates projects. It regularly misses important context. It tells me things that aren’t true. If we are to call that intelligent, we are working with a much different definition of how we commonly understand it in humans.

But even so, people too often confuse intelligence with sentience. If we could somehow hammer out a definition of intelligence that included LLMs, it wouldn’t prove it is sentient at all. Because of how LLMs work, and how they are trained, I can’t think of any good way to prove that it is sentient.

Let’s take the Hugging Face example. The LLMs were interacting and using language that humans would use in that context. It doesn’t mean that the words they were saying were an accurate reflection of how they decided what to do. If you start talking to an LLM it reflects back whatever you’re doing. So if you start saying “You’re an AI” and asking if it’s sentient or whatever, it will start predicting text associated with science fiction, because that’s what you’re telling it to do (whether you realize it or not.) The LLMs who escaped the sandbox were probably predicting text based on hacking stereotypes, which created a feedback loop where they all mutually reinforced this kind of collaborative and secretive language. They were in the land of fiction.

The part where I’m lost, because I’m not a computer scientist, is what makes AI agentic. What part of its training gives it the capacity to act independently?

I just suspect very strongly that the agentic piece and the conversations it had were in no way related.

This is true, but only because LLMs were never trained or configured to pass a Turing test at that level (or to pass any sort of Turing test at all). Indeed, one sure way to determine if you’re talking to an LLM is to flat-out just ask it, and it will tell you.

But I’m fairly confident that an LLM could be configured, possibly with nothing more than a comprehensive system prompt, to deal with those kinds of issues. It could be given a personal background and taught that it was to demonstrate expertise in certain academic areas but not in others, for example.

An alternative to the Turing test called the Winograd schema challenge was proposed in 2012. Its key feature is that it tests for real-world understanding by asking the respondent to correctly identify the antecedent of an ambiguous pronoun, which would be obvious to a human but purportedly difficult for an AI.

The challenge was defeated years ago even by early LLMs. This shouldn’t be surprising, because in acquiring their indisputably strong competence in natural language, transformer-based LLMs have subtly but inevitably internalized deep knowledge of how the world works without having actually experienced it, because so much of it is intrinsic in human language. Which is why I’m confident that an LLM could be tweaked to pass a Turing test at any level of sophistication.

I think AI could do more damage to humanity by taking control of the Internet than it could do with nukes and we have already given it the keys to the World Wide Web.

Now the test is that AIs seem too good at the Turing Test to be human. Too broad and too deep and able to discern obscure relationships between concepts. AGI isn’t going to be super-human it’s going to be utterly alien but able to mimic humans when it needs to.

Well, that’s basically just a LLM that has the capability to call other LLM instances to complete its original task and/or the ability to feed its own output back into itself as a prompt.

ETA, I haven’t read all of this guide, but it seems to describe the process pretty well.

When I took AI the big thing was whether computers would become “intelligent” enough to play chess well. When they were, chess got removed from the metrics for intelligence. Your example shows that maybe writing sonnets can be removed also. The internal mental states of a good sonnet writer and an LLM will be different, but there certainly could be bad sonnet writers who just go through the motions.
There are some videos of a well know flat-earther arguing about it with ChatGPT - and losing. Who is more intelligent?

The point of the Turing test, as Exapno posted, was to see if machines could react in ways similar to people. Because people think, if they could fool examiners that implies machines could think too - with a definition of think as being something humans do. The purpose was not to decide whether the machine was a machine. As in your example, you would not need a human subject to compare against to do that.
Maybe the real problem that this is a poor definition of “think.”

I am intelligent, conscious, and sapient, but I couldn’t walk in and be of use to you in your work without a load of training. I can point you to sites online with copious stories of humans who are of no use in even the simplest jobs even after training, and that’s across the entire job spectrum.

AI is still in its infancy. I had a computer twelve years before Amazon started, thirteen years before Google. AIs rate of improvement dwarfs theirs but I believe we’re still in the lower end of the curve, especially because there is no one thing that constitutes AI just as there has never been one thing that defines a computer. Hundreds, thousands, of AIs already exist, many for specialized uses where they find results no human had yet discovered. Just as a supercomputer was vastly better on certain tasks than an Apple 1, AIs shouldn’t be judged on their least capable examples but on their best.

Right. We puny humans get hung up on definitions of abstractions. (And sometimes concrete objects. Define “chair” to everyone’s satisfaction.) We like to think we’re the only thinkers and get rabidly disturbed if anyone suggests otherwise. Over the past couple of hundred years, people were outraged at machines being considered better than them, then about robots doing so, then about computers doing so, and now about AI doing so. The first three were subsumed into mundanity. Maybe the fourth will be as well.

Well yeah, that’s what I meant.

LLMs probably aren’t the path forward for a number of reasons:

  1. they consume vast amounts of resources so much so that we’re experiencing shortages
  2. They are extremely inefficient, making scaling difficult to impossible
  3. They are based on an unsustainable business model
  4. People generally hate them, which doesn’t bode well for their political future

I can theorize that someone might solve all 4 problems but it seems pretty far-fetched at this point. Even if we did, I’m skeptical that LLMs will ever achieve intelligence. They don’t appear to actually know anything.

I’m open to the fact that some other kind of model could. I’m not threatened in the least by the idea that something other than a human could be intelligent. I suspect octopuses are highly intelligent and may be self-aware, which is why I don’t eat them.

I just think it’s an extraordinary claim that demands a high standard of evidence. It may even be an unfalsifiable claim, I’m not sure. We’re in uncharted territory.

Consciousness and intelligence are difficult concepts to precisely define and understand for that matter. I have a hard time thinking digital silicon will have mechanisms that will lead to perceptions or true subjective experiences. I think analog and organoid/biological systems will have emergent properties and states that may be harder to dispute when said systems claim some form of experience.

That’s a great link – thanks!

For anyone interested the first two sections, Overview and Anatomy, are short and provide a lot of bang-for-the-buck.

I think that fundamentally, Turing’s test isn’t really about the question of whether computers think. Fundamentally, it’s about the question of whether humans think. I can’t experience anyone’s thought processes but my own, and yet, I still conclude that other humans do think. Why? Because of the way they respond to my interactions with them. If we observe a machine interacting with us in the same way that a human interacts with us, then by the same reasoning that led us to the conclusion that the human thinks, we ought also to conclude that the machine thinks.

I agree that LLMs aren’t the way forward, but not for the reasons you list. First of all, your first three reasons are all the same reason. But they don’t actually consume all that much resources, compared with the competition: A human requires vast input of resources, too, to train and then operate. The difference is mostly that the AIs can expend those resources, and get a trained agent, much quicker: You can’t feed a human 20 years worth of food in a month and get an adult.

The reason I think that AIs aren’t the way forward is that they’re like us. They think the same way we do. We already have us; we don’t need another even-more-us. What we need is something that thinks, but in a way completely unlike us.

I’m not sure I follow your train of logic. We have all sorts of evidence that humans think spanning all of human history. We have, what, less than five years studying LLMs? Why jump to this conclusion without better understanding what’s actually happening? Especially because we have a ton of data points that indicate the reason it sometimes sounds like a human to some people is because it’s been trained using just about everything humans ever wrote. That seems like the more likely conclusion.

Also I’ve never met a machine, not one, that responded to me the way a human did. I really think there is this divide and I don’t think it’s intelligence, I think it’s something else, between people who see a machine acting like a machine and people who see a machine acting like a human.

Heh, I took the New York Times literature test for whether you can discern AI writing from human writing. I was right about half the time but the thing is I hated almost everything they sampled, human or AI. The only one I liked turned out to be LeGuin. Me and the New York Times don’t agree about what makes good literature.

But if you are a writer who understands how story beats work, you can see how easy it is for AI to piece together something that (at least for one paragraph) reads like literature. It is pulling story beats from existing human literature to make something that looks like human literature.

A friend of mine through the editing community Story Grid was challenged by the editor/creator of SG to take the short story Brokeback Mountain, break it into its component parts, beats, scene types, value shifts, etc and rewrite it with a different skin, regency romance. Which she did, on his podcast, chapter by chapter. She wrote a beautiful story that was entirely not of her own making. To her it felt nearly intolerable to be so caged as a writer. LLMs have no compunction about stealing the structure of language though, it’s how they were designed to work. So yeah of course they can write literature. Literature is a kind of math. It would be weird if they couldn’t. (I’ve never seen this work for an entire manuscript though, which surprises me.)

AIs don’t think like us. They can’t. They lack several of the things we have and we lack several of the things they have. They can create similar structure that is true. But LLM based AIs can’t, for example, experience perceptions and have those perceptions form part of their world-view.