Melanie Mitchell, a cognitive scientist and computer scientist at the Santa Fe Institute, argues that current methods for evaluating AI intelligence are inadequate and that AI represents a new kind of “alien intelligence.” Her perspective challenges conventional assumptions about how machines process information and whether their outputs reflect genuine understanding or merely sophisticated pattern recognition. This debate has gained urgency as large language models (LLMs) increasingly influence scientific research, policy decisions, and everyday life. The discussion centers on whether these systems truly reason or simply simulate reasoning through statistical techniques. Mitchell highlights the limitations of existing frameworks for assessing machine cognition. Traditional psychological tools used to evaluate intelligence in humans, such as problem-solving tests or behavioral observations, are ill-suited for analyzing the inner workings of AI. These models operate through complex neural networks that lack transparency, making it difficult to determine whether their responses stem from actual comprehension or algorithmic mimicry. This ambiguity has profound implications for how society should regulate and trust AI technologies. The conversation began during a recording session for The Joy of Why, a podcast hosted by Steve Strogatz and Janna Levin of Quanta Magazine. The episode, released on July 23, 2026, marks a follow-up to a previous interview with Mitchell from 2021, shortly after the launch of ChatGPT. At that time, Mitchell already expressed concerns about the gap between human and machine intelligence. Now, with AI systems more advanced than ever, her views have taken on renewed significance. She emphasizes that understanding AI requires adopting methodologies developed for studying intelligence in other species and children, a field known as comparative psychology. Strogatz and Levin noted that the rapid evolution of AI since 2021 has made it challenging to keep up with developments. They acknowledged that even experts like Mitchell must constantly update their perspectives. During the interview, Mitchell discussed how researchers might apply insights from developmental psychology to better interpret AI behavior. For instance, studies of how infants acquire language and solve problems could inform strategies for evaluating whether AI systems engage in similar processes. However, she warned against overestimating the capabilities of current models, citing historical examples such as the mathematical abilities of a horse trained in the early 1900s, which were later revealed to be the result of human intervention rather than innate intelligence. Mitchell proposed six guiding principles for assessing machine cognition, including the importance of transparency, reproducibility, and alignment with human-like reasoning. She argued that until these criteria are met, AI systems remain unreliable as decision-making tools. This stance contrasts with some industry leaders who advocate for greater reliance on AI due to its efficiency and scalability. The tension between these viewpoints reflects broader societal debates about the ethical and practical applications of artificial intelligence. As the conversation unfolded, Mitchell stressed that the challenge lies not only in improving AI performance but also in refining our conceptual framework for understanding intelligence itself. She suggested that the pursuit of creating truly intelligent machines may require a fundamental shift in how we define and measure cognition. This includes rethinking traditional benchmarks and embracing interdisciplinary approaches that integrate insights from neuroscience, philosophy, and computational theory. Such efforts, she believes, are essential for ensuring that AI develops in ways that benefit humanity rather than undermine it. The episode concluded with Mitchell expressing optimism about future progress, while acknowledging the complexity of the task ahead. As AI continues to evolve, the need for rigorous evaluation methods becomes ever more pressing. Whether these systems will eventually achieve a level of reasoning comparable to humans remains uncertain, but the path toward clarity demands both technical innovation and philosophical reflection.
★
Keep the news honest.
ObjectiveNews is reader-funded and ad-free — we show you the bias instead of hiding it. Support independent journalism for €4/month.
Become a Supporter