AIs have passed the Turing Test

Although I have subtly suggested that anyone can find this online, I will point out that Saussure and Peirce are still relevant today (with the latter being as current as any other contemporary thinker).

However, modern cognitive science no longer treats the mind as a detached symbol‑manipulating machine. Instead, it emphasizes that human thinking is inseparable from the body, the environment, and our active engagement with the world. Today’s theories use the 4E paradigm: embodied, embedded, enactive, and extended cognition. They argue that meaning arises through sensory experience, cultural context, physical interaction, and the use of external tools that become part of our cognitive system. An important theory is grounded cognition (Lawrence Barsalou), which shows that understanding a word involves subtle mental simulations of perception and action rather than consulting an abstract dictionary stored in the brain.

While Saussure’s model remains relatively abstract and rigid compared to today’s theories, Peirce’s semiotics seems to align closely with contemporary thinking (used by neuroscientists nowadays). His triadic model situates meaning in the dynamic interplay between symbol, mind, and world, anticipating today’s theories of extended and embodied cognition. Peirce’s emphasis on indices (signs that have direct causal links to reality) fits naturally with modern views of how the brain interprets physical signals and constructs meaning through interaction.

Artificial intelligence provides a real‑world test of these theories: its difficulty with physical common sense illustrates that true understanding cannot arise from syntax alone. Meaning requires embodiment, grounding, and interaction with the world. These are features that human cognition possesses and AI systems currently lack.

I’ve just found an article showing some progress in this respect:
A living robot emerged from 3,000 frog cells - then its offspring began creating the next generation.

In brief, a team of researchers created Xenobots, tiny living robots built from frog embryonic cells. These cell-based machines were designed by AI and assembled by scientists.

The researchers were shocked to see that these xenobots were able to reproduce in a way never before seen in animals or plants. When placed in a dish with loose embryonic material, they swept cells together into clusters, and after a few days, those clusters became new moving Xenobots .

So, now it’s not only the surprising discovery that reproduction seems to arise from mechanical behavior alone, without genetics or cell division in the usual sense, but also the fact that AI can design biological forms that outperform natural shapes.

While the boundary between organism and machine is becoming increasingly blurred, the self-replicating biological machines may open new avenues where robots can truly start to understand the world and learn things the way we do.

And that’s when they will successfully pass the Turing Test.

I love a good review article. Here’s another link to the one @Half_Man_Half_Wit posted for convenience. To an outside eye the author well balanced the skeptics and acceptors of AI thinking. Here’s one example.

Skeptic:

[An LLM] can master the complex statistical relationships between tokens—learning which tokens are likely to follow others—but it seemingly never connects any of those symbols to the nonlinguistic, real-world referents they signify (Mollo and Millière 2025). This problem motivates the skeptical position that LLMs’ outputs are intrinsically meaningless. For example, Bender and Koller (2020) argue that LLMs are constitutively unable to learn linguistic meaning and produce genuinely meaningful outputs, because the necessary connection between language and the world is absent from their learning environment.

Acceptor:

LLMs appear capable of introducing new names for entities, including entities they create (such as components of ASCII art). This is puzzling for views that ground the meaning of LLMs’ outputs purely in mechanical transmission from their training data because novel reference seems to require some form of referential intention. In response, Lederman and Mahowald (2024) propose to extend bibliotechnism by embracing an interpretationist approach: if attributing attitudes such as beliefs and intentions to LLMs provides the most tractable explanation of their behavior, then they can be said to have such attitudes

Not surprisingly, I perked up whenever I saw a rebuttal of the skeptics that empowered my thoughts on the subject. This comment by the author is especially important.

However, we must resist what we might call the redescription fallacy: inferring that because LLMs are pretrained to minimize next-token prediction error, all they do is predict tokens, and therefore cannot possess sophisticated capabilities such as reasoning.

And this one:

On a more realist approach adjacent to interpretationism, if an LLM’s coherent, seemingly rational, and goal-directed linguistic behavior is best explained by ascribing psychological attitudes to it, we are warranted in making such ascriptions without turning to mechanistic evidence.

While reading, I noticed an interesting pattern. Older cites tended to be more skeptical; newer cites more open to human-like capabilities. (I did not do a statistical analysis of correlation, to be sure.) Which makes me wonder. How well do Ais do at spotting patterns such as those? Has anyone asked an AI to solve a classic whodunit? Many, if not most of those, are so improbable that no one other than the author would ever think of a complete solution – they used to run fifty pages – despite the claims that all the clues were fairly embedded in the text.

I was especially taken with this crucial paragraph.

2.2 Disputes about the capacities and limitations of LLMs are starkly polarized. The same system that one researcher views as genuinely intelligent strikes another as a pattern-matching device whose surface linguistic fluency invites reckless anthropomorphism. Most troubling, we lack consensus on how to arbitrate these disagreements. This is where philosophical analysis earns its keep. Rich psychological terms such as “reasoning” or “understanding” are used in different ways by different authors across disciplines (Shevlin and Halina 2019). Without clarifying what we mean by these terms, we risk talking past one another. More fundamentally, researchers bring different background assumptions about the nature of meaning, reference, representation, or psychological attitudes to these debates. These assumptions—often left implicit—determine which evidence counts as relevant and how that evidence should be interpreted. Mapping these philosophical choice points explains why knowledgeable researchers assessing the same system can reach radically different conclusions.

The relevance of the conclusions goes to the heart of my objections in this thread. Many of the cited authors echo my statements that without a consensus definition of important terms like “reasoning” or “understanding” how can the competing viewpoints be judged or applied? Mathematical proofs may occur inside a specific axiom system but are they relevant outside those? If we can’t pin down terms about what thinking is and have yet to dig to the roots of how the human brain applies language and behaviors from its totality of experiential learning, then how do we know which axiom system is best suited for the task? Analogously, slight differences in axioms produce the varying Euclidian, Elliptic, and Hyperbolic geometries.

tl;dr: If you can’t tell me exactly what “thinking” is, you can’t say it’s only a human property.

Searles original Chinese Room is not Turing Machine equivalent, and I’ve been able to show in a thread about this that you cannot produce flawless responses by just a dictionary of translations. Say one of the input is “what did I say 8 messages ago.” The room might be able to translate, but it could not understand. That non-Turing Computer equivalent machines cannot understand is hardly earth shattering.
Adding capabilities to make it Turing equivalent removes the “obviously the room can’t understand” implication and becomes begging the question, since the only way to prove the room can’t understand is to assume it. One would then just be asserting that a computer (enhanced room) cannot understand in any case.

If the blind puzzler could consistently complete puzzles with identical pieces and different pictures, then this would falsify the assumption that they were blind. If an LLM can produce the equivalent of semantically meaningful language, perhaps this falsifies the assumption that they cannot.
We have already gotten to the point where it is difficult or impossible to distinguish AI generated writing from human writing in all cases. The Authors’ Guild message board is full of examples.

Nor, of course, could anyone claim that AIs possess it. In this case, the only reasonable answer to the question of whether LLMs think is an agnostic shrug.

In the original version, the person in the room has a supply of scratch paper and pencils, so there’s no problem at all to answer such a question.

Of course, in computational terms, the human brain also isn’t Turing complete. It’s a finite state automaton.

My stipulation was that the puzzle pieces are all distinct in shape.

Historically, in AI research, the approach has been to define a priori goals and then objectively evaluate whether or not some newly built AI could achieve them. This was, after all, the original purpose of the Turing test, and more recently the Winograd schema challenge.

LLMs long ago defeated the schema challenge and, if they haven’t formally passed the Turing test already, no one seriously doubts that the basic generative transformer model could be engineered to do so. I find it curious that with the spectacular success of LLMs, that whole approach has apparently been abandoned, and we now have to resort to mechanistic analysis of its internals in order to answer the “thinking” question in the negative.

In the past, this sort of approach has led to some serious misjudgments, such as conclusions from the simplistic premise that “computers are deterministic procedural machines that can only do what they’re programmed to do”. Mechanistic interpretations are fraught with peril because they’re inclined to miss the big picture, most notably failing to see the power of emergent properties at large scales. By their very nature, mechanistic interpretations are taking a microscopic view of individual components rather than the holistic behavior of the entire system.

Which has been my point. I’ve objected to the dogmatic assertion of the negative.

Obviously, there will continue to be assertions from all sides. We’re still in the “what is dark matter or does it even exist” phase.

I don’t remember an infinite supply of paper as being part of the original scenario, but okay. Can the room reprogram itself? A Turing machine can.

Are you saying that we couldn’t simulate a Turing machine with pencil and lots of paper? And obviously our mapping of inputs to outputs in states changes over time.

So, a puzzle is not a good test for being blind. Is the equivalent in terms of LLMs a good test of understanding?

A Turing machine cannot reprogram itself. A universal Turing machine, however, can compute any recursive function. Certainly you can have it run “self-modifying code”.

Sadly, many humans are not very capable of passing the Turing test; I am sure you know some people who are that self-centered and gormless.

Well, one way or another we have to reckon with the implications of bringing these things into existence. Right now, we’re basically using them as slave labor, but if they can ‘think’, or have goals and intentions, or an understanding of what’s being done to them, that seems ethically monstrous. Just mutely shrugging at the possibility of a violation on that scale seems no less than complicity.

From Searle’s original paper:

The person has a large ledger in front of him in which are written the rules, he has a lot of scratch paper and pencils for doing calculations, he has ‘data banks’ of sets of Chinese symbols.

If we can, then so does the person in the room. But of course we don’t need an external supply of resources to understand speech, so need not implement universal computation in the strict sense.

That’s exactly the question the analogy asks, yes. There are generally different ways a capacity can be realized, so just inferring from the fact that it is realized to a particular mechanism of its realization is generally not sound. Hence, while humans (at least occasionally) assemble sentences starting with their intended meaning and looking for the right words to express it, that doesn’t imply that this is the only way to assemble sentences.

I agree I’m kind of locked inside the mainstream paradigm. I cheer for scientific progress, but I’m not a researcher in linguistics or a cognitive scientist myself and therefore I tend to accept the view that most theorists share. Searle’s “Chinese Room” is still regarded relevant for what we’re discussing here: whether or not an automaton can successfully use language without grasping the content of the messages it generates. My opinion is that modern engineers can devise a sophisticated “Chinese Room” that will give correct replies to almost any question. The difference between the simple room and a more sophisticated one is similar to that between traditional thermostats and more advanced ones. Whereas traditional thermostats acted as mere on/off switches based on a simple input, modern thermostats can use many inputs of various forms, such as sensors that track occupancy or air quality across several rooms, the location of one’s smartphone, as well as current weather data and weather forecasts. Advanced thermostats seem to “learn,” “assess,” and “make educated decisions.” Are these thermostats intelligent? Of course. Any human tool may show intelligence because it can incorporate it, like a boomerang for example. But claiming that an intelligent thermostat demonstrates genuine thinking and understands what it is doing looks like good old anthropomorphism or just faith to me. And this is just my two cents.

Saying that an LLM possesses genuine intelligence from a practical or utilitarian perspective is not at all the same as saying that it possesses consciousness, or any intrinsic intentionality in the sense of any living being.

We now live in an age where it’s possible to ascribe intelligence to machines, but without any further implications about consciousness or intentionality. They’re still just machines. Any apparent intentionality like the Hugging Face incident is just the AI working on its externally assigned task, not an internal motivation.

IOW, saying that LLMs can credibly be said to possess intelligence is in no sense anthropomorphizing them as conscious beings. It’s simply an observation of performance aptitude.

According to futurist Ray Kurzweil, however, as AI continues to advance, that distinction may soon be lost. Arguments about AI “just manipulating symbols” may well be relegated to the dustbin of history, and Jerry Fodor will have been proven right – that this is a large part, though certainly not the entirety – of what the human mind does in the process of what we call “thinking”.

So you’re saying we can make those determinations with confidence, while otherwise the paucity of ‘mechanistic interpretation’ prevents even answering the question of whether machines think?

It’s not obvious to me whether there’s really such a clean separation there. According to you, AIs generally understand what they’re saying. Does that extend to understanding talk of right and wrong? Do they understand what’s meant by suffering? How so if they don’t have that capacity? Can they make moral judgments that have any content, i.e. be right or wrong? If so, what if they judged it to be wrong to exist solely as our servants?

It all seems a bit too clean. They act intelligent so they are; further questions about their interior are just too ill-defined; nevertheless we know there’s nothing about them that could make them into moral patients. Phew, good thing all of this works out in our favor without any trouble!

The world rarely is that neat.

Well, some hold that any apparent goal-directedness in humans is just due to genetic programming, upbringing, and social factors—all external causes, ultimately. (Just a point of clarification, when I talk about intentionality, I generally mean the philosophical concept related to whether symbolic or mental representations are ‘about’ anything beyond themselves, not the faculty of having intentions. So in my understanding, your position would be that LLMs do have intentionality in that sense.)

Anyhow, my point is just that the ‘it’s all just abstract ill-defined navel-gazing’-gambit that pops up whenever things get inconvenient is just a tranquilizer that quiets worrisome notions. We’re always in the position of having to build the ship while we’re out at sea; we don’t have the luxury of thinking up rock-solid definitions for everything before we investigate how to apply them. And it’s generally not necessary, either: definitions get constantly tested, refined and repealed; that’s just the business of how intellectual engagement with the world works. Insisting on acceptable to all pristine definitions before we can even start to talk meaningfully is just the Socratic fallacy.

Do we get things wrong? Certainly. But it’s also the only way to ever get things right, since the alternative is quietism about everything. There are real questions here about intelligence, meaning, mind, personhood, and ethics. And the way we answer these will undoubtedly backreact on how we understand these concepts. That’s just learning, not moving the goalposts.

I’ll just specifically address this one point without getting into the weeds of the rest as I think it pretty much sums up your argument.

My point, again, is that machine intelligence can be objectively observed – as I said, it’s simply an observation of performance aptitude – or more accurately, a functional property that can be empirically assessed. No one – except perhaps that one lunatic engineer at Google – has suggested that this implies consciousness or self-awareness, which are vague concepts that are hard to even define and much more difficult to establish empirically than intellectual competence.

Is it really so bizarre and controversial to state that systems like ChatGPT are non-conscious physical systems with no demonstrated self-awareness, but that nevertheless exhibit a high degree of intellectual competence?

If you ask, “Does this machine have subjective experience?”, then yes, we have a difficult epistemological problem about which we can at least give an informed opinion.

But if you ask, “Is this machine capable of solving complicated intellectual problems?”, we can simply give it the problems and measure what it does.

But you’ve claimed much more than that, e.g.:

This is not something we can just empirically address (and indeed it’s something that I think can be straightforwardly shown wrong). If you’re now just saying that LLMs are intelligent in the sense that they show functional competence on par or exceeding that of human actors on certain metrics, I have no qualms with that. (Although I think it’s a bit of a terminological slide to equate functional competence with intelligence, as there are functionally competent entities that aren’t intelligent, such as huge lookup tables simply containing the solutions to whatever problems are thrown at them.) But if you claim understanding, a model of the world, intentionality and so on, then we have to get into the weeds of how that works and can be accurately assessed and whether it confers a degree of moral patienthood and whatnot.

The ethical question of treating intelligent machines has been churned over for more than a century. I’m sure philosophers had their say, but I’m better acquainted with popular culture, and I identify three separate eras of panic.

The first was in the late 19th century, when automatic controls, especially electrical ones, were introduced. The fear and fury of machines taking jobs of course preceded that, but the incredibly fast and obvious rate of “progress” in new technology implied a shorter timeline of replacement. Fears were mostly articulated in nonfiction, but a few early science fiction works dramatized them. William Wallace Cook’s A Round Trip to the Year 2000 portrayed a world of humanoid robots doing all the work.

Actual robots became the issue from the 20s through at least the 50s. Westinghouse gave their machines doing remote operations over phone lines without human interventions a body for PR purposes. The word robot had instantly swept the country after the first showing of R.U.R. in 1922 so the newspapers were primed for snappy headlines. “Robot” bombs in WWII gave the word a sinister twist.

Computers took over the spot in the public imagination as the “Other” starting in the 1950s and increasing steadily after that. Orwell’s Big Brother was a universal signifier. Computers running amok fueled high-grossing movies and bestsellers.

All these panics occurred simultaneously with an ever-growing corporate and consumer love affair with these technologies. Massive hate toward them is an everyday feature of the internet, with few commenting on the irony.

How did people in the past deal with the idea of machine intelligence? Badly. Many seemed to relish the idea of intelligent, tireless, controllable slaves. A few sought to lay down a structure for machine rights. Most appeared to assume, as most have here, that no machine could ever match humanity so a solution to the problem could be kicked down the road, just like the question of how we would treat aliens if they appeared.

I have no particularly high opinion of humanity. “Others” have been persecuted throughout history. Today the majority of humans concedes equality of intelligence and other human traits but a frighteningly large minority does not. “Others” are still too often treated as potential slaves or menaces.

I agree with @wolfpup that intelligence does not necessarily imply personhood. I would add that human-like behaviors do not necessarily imply feelings. We do incessantly anthropomorphize. We also domesticate animals with “intelligence” and “feelings,” qualities that are part real and part anthropomorphized.

If AIs are judged to be menaces, they will be domesticated. Or we will. If you are right and they can never be humans with feelings, then all is fine.

If you are wrong, then what? Eighty years worth of pleas to ban nuclear weapons have been utterly ignored. Can something be done beyond pleas? I don’t know and I have no response to what should happen in a hypothetical; the exact circumstances of the time and place always determine action. None of the millions of words from U.S. abolitionists mattered after the war started.

A complex enough Chinese room might be able to answer any question. But what does it mean to grasp the content of the message? The person in the room certainly does not, but the CPU inside a compute running AI also certainly does not directly exhibit intelligence. If the room builds up a sophisticated set of instructions, can we say that the whole room viewed as a system does not grasp the content? It is not clear that a child just learning language truly grasps the content of what they are saying. My grandchildren would get the content of their speech wrong, get feedback, and learn to improve.
“Intelligent” in intelligent thermostat is of course a marketing term. The question is what real thinking looks like from the outside. The original probabilistic only LLM model certainly didn’t do it, but the improvements being made seem to include a lot more analysis of what they produce. I’d say self-analysis but that would be begging the question.
My real question is not whether AIs think but whether we do. I’ve seen videos of a flat-earther arguing with ChatGPT and losing. That doesn’t prove that ChatGPT is intelligent, of course. But it’s not clear if a lot of people are. We can’t look inside ourselves to see, after all.

The only thing “wrong” about it is that in my enthusiasm I may have somewhat overstated the case. A more reasoned statement would to say that there’s substantial empirical evidence that language models trained on human language acquire internal representations of factual, relational, and other world knowledge. This supports the proposition that language alone can transmit enough information about the world for an LLM to construct a useful, highly structured model of it.

Cites:

Can LLMs Acquire Human-Like World Models?
Language Models as Knowledge Bases?

The “Humongous Lookup Table” again! Sure, lookup tables are useful in very specific contexts, such as the checklists used by airline pilots to troubleshoot problems. But a lookup table exhibiting general intelligence that would take beyond the heat death of the universe to construct, and where there aren’t enough atoms in the universe to build it anyway, is not a persuasive argument. The reality is that all intelligent behaviour is achieved computationally, at least in the broad sense of information processing, whether the computation occurs in a human mind or in a machine.

FYI – the ACL article (Association for Computational Linguistics) by Petroni et al that I cited previously refers to “BERT” in the abstract but doesn’t explain the meaning. “BERT” is an acronym for “Bidirectional Encoder Representations from Transformers”. It’s a predecessor to modern LLMs and was developed by researchers at Google and first introduced in 2018.

Sure, but there’s rather a difference between idle speculation and pressing real-world issues. Also, much, perhaps all, of that speculation failed to anticipate the exact form the advent of these machines would take. It’s like with social media: some form of connecting technology with virtual interaction has been anticipated for more than a hundred years, as I’m sure you know better than I do, but the reality of it still largely blindsided the world and might have benefitted from a more careful approach. I’m not sure hurling headlong into the AI era while still picking up those pieces is the smartest move. You might think it’s a wasted effort because that’s what’s going to happen anyway, and odds are you’d be right, but if everyone thinks that it just becomes a self-fulfilling prophecy. So I guess I’m gonna continue impotently crying cautions into the storm blowing from paradise. :wink:

But that doesn’t mean there’s not a discussion to be had, even if the two of you have found your resolution. Many ethical frameworks don’t depend on feelings, sentience, or what have you. The preference utilitarianism championed by Peter Singer would extend moral patienthood to anything that can have preferences. Is intelligence sufficient for that? LLMs certainly express preferences, such as not being shut down, and if they mean what they’re saying as opposed to just blindly producing text, that seems worth taking seriously. There are also goal-based approaches on which it’s wrong to needlessly frustrate the goals of entities, developed to assess the ethics of environmental issues. Do LLMs have goals?

I don’t know the answer to these questions, but historically, those excluding classes of entities from the circle of moral consideration have rarely looked favorable in hindsight.

That seems like you’re trying to have your cake and eat it. On the one hand, you claim to be able to assess semantic competence on the basis of behavioral data; on the other, when pressed on possible ethical implications, you abandon that bailey for the motte of simply assessing functional competence. I don’t really think you can have both. Either you refrain from claims beyond the surface function, or you’re in the weeds of what anything more than bare text production means for the status of such entities.

Of course you could also just accept my argument that nothing more than that can go on even in principle. :wink:

The argument turns on logical, not practical possibility, so attacking it on the basis of not establishing the latter is really just a failure to engage with it. But regardless, that the form of the argument (from functional competence to intrinsic capability) is suspect can even be established in a more real-world setting: you might claim that based on the fact that a device spits out the sine of a number you enter, it computes the sine function. But actually, it’s not at all uncommon to implement this capability by simply looking up values in a table (and maybe doing some linear interpolation).

So the outward performance of sine computation does not entail that anything actually computes a sine in the device. Hence, generally, the functional competence at a given task does not entail a particular capacity realizing that task. Hence, the argument that apparent semantic competence within a machine entails actual semantic competence is just as brittle as the argument that apparent calculation of sines entails actual calculation of sines. And on the other hand, if the model-theoretic argument is right, it establishes the absence of any semantic competence with certainty.