When an LLM answers a question, is it reasoning like humans, or just producing text that looks like reasoning? The distinction isn’t just philosophical, this determines what we can trust AI to do, how closely we need to supervise it, and ultimately what its real-world impact will turn out to be.

Melanie Mitchell (opens a new tab) at the Santa Fe Institute argues that we lack adequate methods for measuring machine cognition, and that AI is a form of “alien intelligence” that operates through non-human cognitive mechanisms. In this episode of The Joy of Why, Mitchell tells Steven Strogatz how methods that psychologists use to study cognition in other kinds of “alien intelligence” — babies and animals — can be adapted to probe AI, and she lays out six principles for better assessing machine cognition. Their conversation ranges from the challenge of interpreting what’s happening inside these systems, to recent AI-assisted breakthroughs in mathematics, to why a math-performing horse from the early 1900s offers a cautionary tale for how we assess intelligence.

To read more, click here.