That information is only there if you already assume that the sentence is being understood, which is question-begging. The problem is exactly whether, and how, one can build understanding just from what that sentence is without assuming that it is already understood: a sequence of tokens that don’t in and of themselves carry any meaning.
And yet, just from those supposed “placeholder relationships”, LLMs can answer many questions about the real world and solve complex problems that even most humans cannot. How do you suppose they accomplish this magic?
ETA: I don’t intend that to be a snarky question. I think the answer is in the rich set of semantic relationships that are established within its neural network, where at a sufficiently large scale something that can fairly be regarded as “real” understanding begins to flourish – most notably, intrinsic information that cannot be traced to explicit token relationships.
Semantic information is embedded within language, and it is formalized through grammar and syntactic rules. This form expresses, organizes and transmits concepts generated at non-linguistic level. Although AI systems master the use of language and language integrates semantic information, an AI lacks the type of conceptualization and understanding that may allow it to achieve true cognition.
Well, I don’t really need to know how something works in order to understand how it doesn’t—I don’t need to know the workings of the magician’s illusion to know that nothing supernatural is going on. That what on the surface seems to be the case actually isn’t. And of course, it’s always a bad idea that just because I can’t imagine any other way for something to work, there isn’t one.
Regardless, I have given at least a rudimentary description of how I think the trick works:
Knowing the relation one token stands in, and knowing the relations other tokens stand in, allows at least probabilistic inferences about further relations these tokens stand in. More generally, entailment is syntactical: one need not know what A and B stand for in order to know that ‘if A\rightarrow B and \neg B, then \neg A’. On a formal mathematical logical level, knowing the axioms of a system (which is equivalent to fully knowing its structure in the relational sense) allows you to derive all sorts of theorems, but it doesn’t allow you to pin down a model of the system uniquely.
Brian Merchant’s recent book, Blood in the Machine: The Origins of the Rebellion against Big Tech, has much to say about Luddites past and present and of tech past and present and future.
One point Luddites made was it is not the technology that poses the problem, it is who controls it and for what purpose. The original Luddites would have loved the new machines if they were used to cut their work day in half for the same or more pay.
Instead, the machines were used by their owners, who were responsible to no one but themselves and their shareholders, to make half the workforce unemployed and work those remaining harder and longer in the quest for private profit. The problem of AI is a problem of ethics and politics and economic systems–who is in charge–as well as one of “technology.”
Finally, I keep misreading this thread’s heading as “I have passed the Turing Test,” and wonder how I’d score on it.
My understanding is it took roughly 60-100 years for the benefits of the industrial revolution to start meaningfully trickling down to the working class, and a lot of that was due to political changes forcing the benefits to be shared.
A world where evil greedy psychopaths like Elon Musk and Peter Thiel control AI and use it to implement wealth inequality, global fascism and mass unemployment, as well as environmental destruction, is terrifying for a lot of people. Then you have the fears of runaway AI that can easily outsmart its handlers.
I have several objections to this line of argument. First of all, it’s doubtful that major entities are spending tens if not hundreds of billions of dollars to develop and commercialize a mere illusion or “trick” as you put it. Said another way, if LLMs exhibit intelligent behaviour and solve complex problems in the vast majority of cases, then they should be regarded as intelligent if only from a purely utilitarian perspective. The claim of lack of “true understanding” becomes a question of metaphysical semantics.
Further, I have some issues with your proof example. You’ve constructed it with the explicit stipulation that the symbols have no meaning to the system. But that’s not what we’re debating. The real question is “Does the actual structure learned by an LLM contain representations that correspond systematically to semantic properties of things in the world?” Your example doesn’t illuminate that question.
The structure learned by an LLM is not merely an arbitrary relational structure among meaningless symbols. It’s a learned statistical representation of a corpus whose structure is systematically related to the structure of the world. So whether the resulting internal representations capture aspects of that world is an empirical question and not one that can be resolved by philosophical abstractions.
The mere fact that the computation can be described syntactically does not establish that the representations are semantically empty. IOW, your statement that “it needs no interpretation to do so” is true in this example, but it doesn’t establish that no interpretation is actually present in an LLM. Saying the LLM “is just manipulating symbols” isn’t an empirical conclusion, but a philosophical interpretation.
This all reminds me of the old, old argument from the early days of AI. It went something like this: Computers are procedural things that execute instructions from a stored program in a rote, deterministic fashion. They can only do what they’re programmed to do. Therefore, computers can never exhibit intelligent behaviour, or solve problems requiring intelligence.
The first sentence is true, the second is true but profoundly misleading, and the conclusion in the third is manifestly false.
I should add that while I fundamentally disagree with your position here, I always appreciate your thoughtful arguments and insights.
There’s something about the “just manipulating symbols” argument that’s oddly familiar, and it just occurred to me what it is. It’s from John Searle’s “Chinese Room” argument! Which I always thought was deeply flawed, for reasons that I think have already been extensively discussed in other threads. I think it’s very relevant to discussions about “understanding” in LLMs. The moral of the story, in my humble mind untrained in philosophy, is that you cannot reach sensible conclusions about “understanding” by examining individual underlying components of the system in question – you have to examine the system in its holistic entirety.
This is not necessarily an argument that the behaviorist approach is the only one that makes sense. It is, however, an argument that you cannot infer the absence of a system-level property merely from the fact that its individual components don’t appear to possess that property. It’s my old argument about emergent properties all over again. Individual water molecules aren’t wet, individual neurons in the human brain don’t understand English. But the totality of the organized system does.
I’m really trying not to sound snarky here, but are you actually arguing that we should believe AI has actual understanding because big corporations wouldn’t deceive people?
Sure, because that’s what you asked me to do. I’ve given an argument to the effect that, given how LLMs work, they can’t actually possess any understanding, and you asked how they then could answer complex questions about the real world, which is what the example aims to show.
Again, that’s what the model-theoretic argument does.
Exactly. And from there it follows that an LLM can’t possess any understanding of the world. Again, that’s just Newman’s problem: Russell’s original position was exactly that all we know is a structure that’s systematically related to the real world, to which Newman then responded that if that were true, then all we’d really know about the real world is just the minimum number of objects within it, and nothing else. For Russell, this was what has been described as a ‘classic Homer Simpson ‘Doh!’-moment’.
No, but the fact that from merely a set of relations, we can construct a great number of models of those relations—that they don’t fix meanings, in other words—does.
I’ve not claimed it’s an empirical conclusion; it’s a logical one.
That’s not the argument I’m making though. Searle’s argument was that from merely manipulating symbols, an agent can’t infer their meanings. That’s true as far as it goes, but is vulnerable (as you say) to the ‘systems reply’. I’m saying that what goes into the system—training data in the form of ordered sequences of symbols/tokens—does not contain any information from which it is possible for any sort of process to construct a unique mapping of these symbols to the real world. There’s no vulnerability to the systems reply here: nothing can conjure up from thin air what isn’t there in the first place.
Just to address this one point, that’s not my argument at all. My argument is clarified in the very next sentence, with the most relevant part emphasized here:
I’m obviously not arguing that we should have blind faith in AI tech companies, just as I’m sure you’re not arguing that they’re collectively investing hundreds of billions of dollars in an outright scam. My argument is that LLMs have tremendous utility as engines of knowledge and problem-solving. If there’s an underhanded scheme behind the massive AI tech investments, it’s probably the expectation of massive displacement of knowledge workers, especially at lower levels, and hence tremendous profits all around.
Sure. My point is just that the inference from ‘they show human-like competence at language production’ to ‘they must have a a deep if imperfect human-like understanding of semantics in real-world contexts’ is brittle at best, akin to the inference from ‘the blind puzzler can complete puzzles in a skilled fashion’ to ‘the blind puzzler must be sighted after all’. Same surface capability, different realization. The model-theoretic argument then aims to establish that LLMs can’t actually have such human-like understanding, because what they have to work with simply doesn’t suffice for that.
I agree AI is a great achievement despite its limitations, of which specialists are very well aware. It isn’t impossible that future machines should incorporate biological material, which may allow them to fully replicate human performance. Or maybe quantum computers will be able to develop capabilities we can’t even envision.
But right now, the most important names I can think of are Ferdinand de Saussure, Charles Sanders Peirce, and John Searle.
Ferdinand de Saussure described the linguistic sign as a two‑part psychological unit made of a signifier and a signified. The signified is the mental concept we hold in our minds, while the signifier is the internal echo of a word we hear when we speak or read silently. These two aspects form an inseparable whole, like two sides of a sheet of paper. We can apply this model to single words, phrases, idioms, and full sentences.
When this framework is applied to artificial intelligence, we can see that AI does not possess linguistic signs in the human sense because it lacks the psychological component required for signification. AI operates entirely on value. It does not know what a “scarecrow” is; it only computes how the symbol relates to other symbols in vast statistical patterns. Meaning arises only when a human interprets the output. The AI produces signifiers, but the signifieds live only in the reader’s mind.
Peirce’s triadic model clarifies this further. An AI can manipulate the representamen (the symbol itself) and, if equipped with sensors, can even interact with the external object the symbol refers to. But the interpretant (the internal understanding that connects symbol to world) is absent (or must be redefined as a computational state rather than a subjective experience).
Searle’s “Chinese Room” thought experiment illustrates this distinction probably more clearly. A person who does not understand Chinese can still produce flawless Chinese responses by following syntactic rules, convincing outsiders that genuine understanding is taking place. Yet inside the room, there is only mechanical symbol manipulation. Searle argued that this demonstrates the difference between syntax and semantics: a system can simulate understanding without possessing it. Large language models operate in the same way, generating statistically appropriate symbol sequences without any internal grasp of their meaning.
Has the field of linguistics and cognitive sciences in general moved on from structuralism as well as post-structuralism and is now in an unnamed phase that could be nicknamed “post-post-structuralism”?
This clearly isn’t an area of strength for me, but it seems like we are analyzing AI with constructs from the 60’s and 70’s.
It was an honest question. I admitted this is not my field.
In Physics we augment Einstein’s work with other subsequent physicists, particularly in areas not as heavily understood in Einstein’s time.
Cognitive sciences has also progressed rapidly. Does it make sense to focus on constructs from Ferdinand de Saussure and exclude Chomsky, Derrida, and later?
I’m only here to learn more about this topic. I started to read about structuralism to understand the points being made. I discovered that structuralism is two eras old and not the latest thinking in cognitive sciences.
Just as Einstein is still being proven right by current observations, the linguists and philosophers I mentioned are being proven right by the current studies cited earlier in this thread. (Anyone can easily find them on the Internet with the help of… AI.)
What does it mean to by proven right in this field? My understanding is that the two‑part psychological unit and the triad are constructs to theorize, explore, and better understand cognition. They aren’t verifiable through experimentation.
Why is structuraliam the right framework to evaluate LLM-based AI? Are there no newer and better frameworks in cognitive science?