In 1998, Robert Horn, Jeff Yoshimi, Mark Deering, and I published seven large-format argument maps under the title Can Computers Think? The History and Status of the Debate. I authored Maps 1, 4, and 6: the broad map of the debate, the map of the Chinese Room argument, and the map addressing whether computers must be conscious to think.
The maps attempted something that ordinary prose handles poorly. The debate over machine intelligence was not one argument with two sides. It was a network of questions about computation, representation, consciousness, creativity, free will, formal systems, causal powers, language, embodiment, and the relationship between performance and understanding. An answer in one part of the network often depended on assumptions buried in another.
Argument mapping made that structure visible. It forced us to distinguish a claim from the evidence offered for it, an objection from a restatement, and an unresolved stopping point from a conclusion. It also made intellectual disagreement look less like a contest of personalities and more like a navigable structure.
Nearly three decades later, the technology is profoundly different. The discipline required to evaluate it is not.
What changed after 1998
The systems represented in the original maps belonged largely to an era of symbolic AI, early connectionism, cognitive modeling, and philosophical thought experiments. Today we encounter machine intelligence through foundation models trained at enormous scale, multimodal systems, tool-using agents, and language models that produce fluent, context-sensitive responses across a vast range of topics.
That change matters. Contemporary systems do things that many participants in the earlier debate did not expect to see so soon. They can write usable code, translate among languages, summarize technical literatures, describe images, generate design alternatives, and sustain long conversations. In professional work, they can compress hours of search, drafting, comparison, and reformulation into minutes.
The existence of these capabilities should rule out a lazy skepticism that treats every apparent advance as a trick. Large language models provide genuine leverage. The interesting question is where that leverage comes from and how far it extends.
What did not change
Fluent behavior does not settle every question about the mechanism producing it. A system can be useful without reproducing human cognition. It can solve human-relevant problems without possessing the same understanding, causal model, background knowledge, or conscious experience that a person brings to the task.
This was one of the central lessons of the original debate. The Turing Test asks what can be inferred from behavior. The Chinese Room asks whether syntactic success is sufficient for semantic understanding. Arguments about physical symbol systems, connectionism, consciousness, and mathematical limits ask what kinds of architecture could support thought—and whether there are principled limits to formal or computational accounts.
Those questions are not museum pieces. They return whenever a model generates an impressive answer and we move too quickly from “the answer is useful” to “the system therefore reasons, understands, or judges as we do.”
From thinking to strategic judgment
My 2026 paper Mean Articulation Machines moves this problem into strategy. It does not ask whether LLMs are generally intelligent. It asks which activities in strategic cognition match what current systems do well and which depend on capabilities their architecture does not reliably supply.
The paper describes a continuum. At one end are tasks that benefit from broad association with documented knowledge: search, aggregation, synthesis, paraphrase, comparison, professional drafting, and the articulation of established frameworks. LLMs are at their strongest here.
At the other end are tasks requiring sustained causal explanation, counter-consensus reasoning, deductive consistency, long-horizon planning, genuinely novel theory construction, or access to tacit and undocumented knowledge. The further a task moves toward this end, the more dangerous it becomes to equate fluent output with sound strategic judgment.
This is why “Can an AI do strategy?” is the wrong level of question. Strategy is not a single task. It includes gathering information, mapping a landscape, identifying assumptions, constructing explanations, making predictions, imagining alternatives, committing resources, coordinating people, and acting under uncertainty. Different parts of that work require different capabilities.
The continuing value of maps
The original maps offered no single verdict on whether computers can think. They offered something more useful: an organized view of the arguments needed to make a verdict responsible.
I see Mean Articulation Machines in the same spirit. Its continuum of strategic tasks is another kind of map. It is meant to help scholars and practitioners locate a task before selecting a tool, distinguish impressive articulation from the particular form of judgment required, and revise the boundary as systems improve.
The technologies will continue to change. Our categories and tests must change with them. But careful evaluation will still require the habits that argument mapping made unavoidable: separate claims from evidence, make assumptions visible, follow objections to their stopping points, and resist substituting a vivid performance for an explanation of what produced it.