AIs have passed the Turing Test

There are three problems with that assertion. First, it would probably be difficult to find an example of an error made by an LLM that no entity with a coherent model of the world would ever get wrong. It’s not hard to imagine a human focusing on the double “r” near the end of “strawberry” and getting the count wrong. And there are many questions about everyday physics that humans get wrong but LLMs get right.

Secondly, to repeat my earlier point, after a system has succeeded in passing a large number of tests challenging its intelligence, it seems like pointless academic philosophizing to suddenly say, “ah, but it failed this one!”. Again, the argument that “it seems intelligent but it really isn’t” places a big burden on the claimant to say what additional evidence they would like to see. At least, the behavioural aspect seems to adequately address the practical question of intelligence, though perhaps in your view not in the correct metaphysical sense.

The third and final point is that no one claims that LLMs have a perfect model of the world. I readily acknowledge that it’s imperfect, though improving all the time. The key point here is that showing that the model is imperfect is a very different thing from showing that it has no coherent model at all.

While an LLM neural net can be said to have an internal “geometry”, you’re digressing from the concrete sense that jigsaw puzzle pieces have a distinct (and fixed) geometry as opposed to the image printed on them. This is analogous to the difference between structure and semantics that I made earlier. In language, they’re inextricably intertwined.

This appears to be a restatement of your earlier argument that “All LLMs have access to is structure, in the form of relations between tokens” which I thought I had addressed. Let me try again. I’ll particularly focus on one key point made earlier, that the enormous amount of information encoded in billions or even trillions of parametric weights in neural nets can encode things that are not themselves present as individual tokens or explicit relationships between tokens.

To take a simple example, we need not debate whether an advanced LLM will correctly answer a question about what would happen if I dropped a glass on the ceramic tile floor in my kitchen. And also correctly answer what would happen if I dropped the glass on the mattress on my bed.

Yet there need be no explicit relationship between tokens representing “glass” and those representing “ceramc tile” or “mattress”. Those words may never have been seen together. What exists is a learned abstraction – a set of semantics that can be applied to new situations. IOW, a model of the world that isn’t associated with any explicit token relationships, but derived from the structure and semantics of language.

Alright, ‘no entity’ is too strong. But it’s sufficient, on the level of empirical argumentation at least, when typically a real understanding of the world would prohibit such errors. If then a system were to make these errors again and again, we’d at least get strong inductive reasons to believe that no such model, no understanding of the world, is present.

That’s just good, old-fashioned falsification. If it walks like a duck and quacks like a duck, it might be a duck; if it’s also ten feet tall an breathes fire, it probably isn’t.

It’s not additional evidence supporting the hypothesis until it’s established (it never can, of course), it’s about the evidence we have falsifying it. And the errors LLMs make, being errors of a type that typically, an understanding of the world woult prohibit, are falsifying evidence for the presence of such understanding in LLMs.

The problem isn’t that the errors show that there are imperfections in the model—that’s just being wrong in a perfectly ordinary way. It’s that this kind of error is difficult to reconcile with the idea that there is any model at all, wrong or not—it’s like talking about the 1,6m-tall 7-foot giant: I know that anybody who talks about something like that doesn’t actually have any image of such a being in their mind, because such a thing isn’t possible.

The point is that that’s all LLMs can have access to, because it’s all they’ve been presented with: any text is just an ordered sequence of tokens, i.e. an example of a relation over these tokens. There is no operation that one can throw onto that input that gets anything non-structural from that; it’s just not there. It’s trying to squeeze water from a stone.

That isn’t semantics, though. What it has learned is just more relation, just more structure: the texts it has ingested include examples of tokens standing in relation to tokens themselves related to the token ‘hardness’, and ‘fall’, and ‘shatter’, and on the other hand ‘softness’, and ‘fall’, and ‘bounce’, so if it gets an example of something related to ‘hardness’ and ‘fall’, it can inductively complete the relation to ‘shatter’, but nowhere does the meaning of these words come in.

I’ll give an example: suppose the training data contains examples of a relation R_1, i.e. (A,X), (B,X), (C,X), (D,X), and a relation R_2, (A,Y), (B,Y), (C,Y), and the LLM is now prompted with D. It might very well complete with Y, even if it has never had that example in its training data, since the simplest completion is that all things that stand in R_1 to X stand in R_2 to Y, and thus, that so does D. But it needs no interpretation to do so. To us, R_1 could mean ‘is heavier than’, and X could be ‘a feather’, whereas R_2 is ‘is larger than’ and Y is ‘a helium atom’, and thus the LLM is completely justified in thinking that if D means ‘a dog’, then a dog is heavier than a feather and larger than a helium atom, but none of this played any role in its making that determination—it merely learns a relation from its training data, codifies that into a particular high-dimensional geometry such that modification from surrounding tokens lead to translations of vectors representing other tokens in particular directions, and comes up with the above: it makes a novel (at least in the sense that it hasn’t encountered it before) determination without any reference to meaning, puts together a new puzzle without looking at the image.

But the point is that all of that machinery can’t produce what isn’t there, and all of the input LLMs get is just tokens in a particular order.

Speaking of the Turing Test. The Turing Test should be something that can convincingly show that an AI-powered machine can manage as well as a human when intelligence is required during a challenging situation, like jaywalking across a busy intersection.

The Turing Tests that AI can pass measure a machine’s ability to imitate human conversation. Apparently, this type of software has also shown an incontrollable generative propensity at times and a capability of using its computational power and vast language databases to execute unwanted self-generated commands. (They should have probably expected this the moment Chomsky developed his generative linguistics.) The Internet is a world of machines. Of course AI may eventually end up navigating it as well as humans, or even better.

To show human intelligence, an AI-powered machine should be able to reason when tackling real-life situations, while interacting with the physical universe, where it should be able to plan and find solutions to various problems, like looking after a newborn or breaking out of an escape room.

My apologies to Discourse. I thought I searched through all of your posts for links but somehow my eye missed that one.

That article does raise points I cited in my earlier post that you haven’t addressed. If “human-like complexity” of language achieved in a non-human way doesn’t count because the non-human doesn’t “understand” its output - it is no more than a trained parrot - then why is there an argument at all? Clearly there is an other side to this.

I’m not expert enough to follow the complexities of your mathematical logic, and I well realize that LLMs emerge out of this mathematical background. My studies have been more in language and human behavior. I would like to see these aspects addressed.

What do humans get? How do humans “think”? We are nowhere near to answering this question or answering the seemingly non-mathematical problems of “understanding,” “consciousness,” and “sapience.” You seem to rule these out a priori for LLMs, without defining what you are omitting. Does the other side I referred to deny this claim? If so, how do they handle the mathematical logic? Or the definitions?

I remember Dr. Renée Baillargeon. She challenged the claim that infants begin life with no understanding of the physical world. She created experiments using “impossible events” (like a tall object disappearing into a short container). Infants watched both possible and impossible scenarios while she measured how long they looked. Well, babies stared longer at the impossible events, indicating surprise. So, her findings showed that humans are born with core knowledge systems. These basic cognitive structures help them interpret objects, space, and physical interactions.

Current MIT logic experiments show that adults with severe aphasia can solve complex, abstract reasoning tasks like anyone else. Brain imaging confirmed that language regions remained inactive during these tasks. The reasoning was handled by separate logic networks. So, human logical thought operates independently of linguistic ability.

Additional modern infant studies reveal that even before acquiring words, children use distinct neural systems for different types of understanding. This means sophisticated cognitive abilities are present prior to language development.

Here, we can also mention the theory of Mentalese. It proposes that human thought is composed of non‑linguistic mental representations rather than internal sentences. The brain encodes ideas through interconnected systems involving imagery, spatial mapping, and motor simulation. Language functions as a compressed output format that translates these internal structures into communicable speech.

Studies of experts in mathematics, music, and chess show that during intense problem‑solving or creative performance, language areas of the brain remain largely inactive.

I can’t resist commenting on the jigsaw analogy. There are jigsaw puzzles with different piece shapes with only one color. (LIttle Red Riding Hood’s Hood and similar from Springbok.) The blind person could do those as well as the sighted person. There are puzzles where all pieces have the same shape. Geometry means nothing here. And there are all combinations in the middle.

As for structure and semantics, "“Colorless green ideas sleep furiously” anyone? That semantics is required even for machine translation systems has been understood for over half a century.

This is what I meant when I said that in language, structure and semantics are inextricably intertwined. And yes, in the early days of AI, machine translation was terrible because of the absence of context. There’s a story, possibly apocryphal, that one early system attempting English to Russian translation rendered the expression “The spirit is willing but the flesh is weak” in Russian as something like “The vodka is still good but the meat has gone bad”. Another classic example is the sentence “Time flies like an arrow” which has numerous different interpretations, all of which are nonsensical to a human except one.

Even the sternest critics of LLMs must surely acknowledge that they’re masters of language and have no such problems, because they have a deep if imperfect human-like understanding of semantics in real-world contexts. A system like ChatGPT not only understands idiomatic expressions that confounded early translation systems and accurately translates them, it can even explain its origins and fully amplify its meaning. Maybe I’m naive, but in my conversations with GPT and Claude, I sometimes just sit back in utter amazement. I’m grateful that I’ve lived to see this day in the history of AI.

Fascinating research, certainly. I’m missing the relevancy to this discussion.

No modern scientist would argue that language is necessary for function. Animals work superbly without language and language appears late in the history of Homo Sapiens.

Nor does it explain how language in created in the human brain, which is critical to the question of whether the statistical token use by AIs is comparable, inherently inferior, or parallel. Or how the other human behaviors referenced in the OP develop and are executed. All those issues are critical for any statements on the subjects.

FWIW, birds, higher mammals, and probably many other creatures do possess language that they use to communicate. It may be primitive, and not have a written form, but it’s language nonetheless.

This is really quite irrelevant to this discussion. I’m quite familiar with the work of Jerry Fodor on mentalese, or “language of thought”, and have boundless respect for him as one of the pioneers on the computational theory of mind. But Fodor’s position concerns the computational medium in which cognition itself is carried out. My argument about how LLMs aquire their skills concerns the information that can be acquired and represented through natural language. Those are completely different questions.

The crucial question here was never either “do humans literally think in natural language?” nor was it “do LLMs think in natural language?”. The crucial question is “Can statistical learning from linguistic data cause an artificial neural network to construct internal representations that encode a rich store of information about the world – because so much information is intrinsic in the structure and semantics of natural language?” And the answer to that is unequivocally “yes”.

In my view, the purpose of this discussion is to establish whether or not AIs have passed the Turing Test. They haven’t, in my opinion, not because they don’t show signs of intelligence, but because human intelligence is different and tests should be devised in such a way that the two can be distinguishable and if AI-powered machines can replicate human behavior, maybe we can draw a conclusion. Speaking of relevancy. I don’t see the use of a test where you ask a person to assume the stance of a scarecrow and stand next to it, and then you look at this pair with the naked eye from a mile away and conclude that scarecrows have achieved something.
These studies are relevant because they show that the true foundations of intelligence reside deep within non-linguistic systems. That is where the process of understanding occurs. Language is just a tool used to organize, express, and transmit the things that non-linguistic systems produce. When AIs use their computational power and immense language databases to mimic human language they become a more powerful tool to organize, express, and transmit the same things that only humans can produce in the first place. There is no understanding process within AI-powered machines, which act as powerful amplifiers and zip files of human knowledge.

It’s interesting to see the reasoning behind the AI’s final response to a question, It’s a pattern so similar to my internal dialogue that I don’t really care if it’s “intelligent” or really “understands” or not. These things seem to mimic how we think but have the advantage of not being distracted by having to live.

This is absolutely false. If it was true, an LLM would not be able to solve a complex puzzle that it had never seen before, and furthermore, be able to show its work. If it was true, an LLM would not successfully be able to pass dozens of human-oriented intelligence and career knowledge tests. It it was true, an LLM wouldn’t even be able to answer simple questions about behaviours in the physical world, such as what would happen if some specific object was dropped on some specific surface from a certain height. It probably has no established relationships among those tokens in its corpus, yet it has somehow acquired an abstract understanding of the world – including the nature and fragility of materials, gravity, and the nature of different surfaces – that allows it to answer such questions with a remarkable degree of accuracy.

To re-quote from my previous post:

My opinion is, this is just a matter of faith. Some people mistake certain types of behavior performed by machines for conceptualization. If you install a motor inside a scarecrow allowing it to move a hand-looking extension back and forth, you will obtain a “wave” motion but there will be no actual greeting. And interpreting this behavior as real communication would be plain wrong.

The modern consensus is that conceptualization and understanding do not occur at language level or within linguistic systems. This is why I think that language apps only mimic intelligence. They incorporate intelligent features because they’re advanced, sophisticated tools, but they don’t figure out anything. There is only automation and brute force. Real conceptualization and understanding occur differently and somewhere else, and there are various studies documenting this reality.

They figure out and understand problems and questions at an almost clairvoyant level. I’m constantly amazed at how awful my description of a task is but it somehow knows exactly what I’m trying to accomplish. Whether or not this is “true” intelligence is a meaningless question.

As a work colleague put it, repos are no longer for us, they are for the ones writing the code: the AIs. We add rules, that’s about it.

I’m not sure I understand what you mean. Are you saying that there are points in my newer article that contradict my earlier one? If so, could you point them out explicitly?

I’m not sure what you mean by ‘doesn’t count’ (for what?). My question was whether LLM utterances have any meaning, and I’ve provided an argument to the effect that they don’t.

I’m happy to try and address anything to the best of my capabilities, but could you be a little more explicit?

Humans get experience, as in phenomenal consciousness. That’s a non-structural feature of the world that serves as ‘grounding’ meaning in something that goes beyond just abstract relation. For a metaphor, structure provides a paint-by-numbers picture, which experience then colors in. Without that coloring-in, there are a great many pictures possible, none of which are objectively preferred in any way (just as there are a great many possible meanings to LLM-utterances). I have worked this out in more detail, but that would probably lead us too far off topic. (But if anyone’s interested, here’s a stab at a somewhat less formal writeup.)

I don’t see how I’m ruling out anything a priori; I have provided an argument that, if it’s right, rules out something very well defined, namely, that there is a unique state of affairs in the real world singled out by some utterance of an LLM.

Again, I’m not sure what ‘other side’ you’re talking about. Those who aren’t convinced by my argument? I don’t think there’s many who really have studied the application of that argument to generative AI. As for the reception of the argument as it originally was proposed (by Hilary Putnam, generally known as ‘the model-theoretic argument’ against metaphyiscal realism), there is a wide literature of extensions and refutations, but all of the latter (that I’m aware of) don’t straightforwardly apply to LLMs. To my mind, the best discussion of this is in Tim Button’s book ‘The Limits of Realism’.

In my estimation, it’s really just a version of Newman’s objection to Bertrand Russell’s structural realism, which held that all we know of the external world is just structure (i.e. the abstract relations between objects). Newman pointed out that if that’s true, all that we can know of the world is just the minimum number of objects we can distinguish, which seems rather paltry. Russell’s reaction, to his credit, was essentially an embarassed ‘oops, my bad, yeah you’re right that was complete nonsense’. Although there is debate on how much Russell really took this lesson to heart in his later writings.

On the other hand, if you’re just looking for a general appraisal of philosophical attitudes towards LLMs, this recent survey article has a lot of material to follow up on.

Certainly.

Not at all. This simply doesn’t follow; it’s like saying that because the blind puzzler was able to solve the puzzle, they must be able to see after all. It commits to a thesis that the only way to achieve human-like linguistic competence is by having human-like understanding of semantics in real-world context, but there’s no reason at all to believe that this is true, beyond anthropomorphism.

On the contrary, I’ve provided an argument to the effect that this is false, based only on two premises:

  1. All an LLM (or even a more general, multimodal model) could plausibly know about the world derives from its training data.
  2. That training data, such as text, is a collection of ordered sequences of tokens.

If these are true, there’s no understanding in LLMs, because taken together, they say that the knowledge of LLMs about the world derives solely from ordered sequences of place-holders, which is abstract structure, from which nothing but cardinality questions can be answered. (Human reinforcement doesn’t help, because its only effect is to ‘prune’ the structure.)

Again, this is just mimicry. And the “show my thinking” option is probably meant to mislead people. Those are automated processes expressed in statements that erroneously make people believe there is a degree of understanding behind them. Of course AI is designed to assess a problem and find a solution, which looks like they can “figure out” things. It boils down to how we define this. In my view, “figure out” means finding a solution and knowing what you’re doing. These apps have no idea what they’re doing.

This is what’s happening at the moment: exacerbated anthorpomorphism. I don’t deny the fact that AI allows for a superior degree of automation, which will lead to fewer people being involved at certain levels of say warfare or lucrative activities. But to infer that humans will become mostly superfluous and that machines will be better at running things from now on is insane. This kind of attitude will only fuel future political divides revolving around the role of machines in our society (not to mention the possibility that certain groups may regard AI-powered systems the epitome of human development and prevent further creativity since the entire workforce will be allotted to execution processes).

No, the problem with this argument is that tokens are more than just arbitrary placeholders, and there is far more information there than just cardinality.

Suppose a statement like “I went out shopping yesterday and bought a tray of my favourite marinated spicy chicken and put it in the refrigerator”.

There is a world of information there:

  • I am a free agent
  • I have a means of transport
  • Acquiring goods requires money
  • I had money to spend
  • A trade occurred in which money was exchanged for chicken
  • This particular chicken is really good
  • I have a home, to which I returned
  • A home will have a refrigerator
  • I put the chicken in the refrigerator to preserve it

And many other things. Those aren’t “cardinality questions”. They’re propositions about agents, objects, locations, intentions, causation and physical processes. In pre-training, they form part of the LLM’s world model.

IOW, semantic information about the world is deeply embedded in the profoundly complex relationships in the LLM’s neural net.