Pathway has introduced a new post-transformer AI model, Baby Dragon Hatchling (BDH), which overcomes the limitations of transformers by enabling generalization over time.
Pathway, the data company building live AI that thinks in real time like humans do, has announced Baby Dragon Hatchling (BDH), a new ‘post-transformer’ architecture that addresses one of the most significant barriers to autonomous artificial intelligence (AI): the inability to generalize over time.
Generalization over time – the capacity to sustain reasoning, learn from experience and make predictions based on new information – is a fundamental property of human intelligence. Today’s transformer-based models, by contrast, are powerful but static: they excel at pattern-matching past data but have limited ability to extend their reasoning into new contexts.
In its new paper – The Missing Link Between the Transformer and Models of the Brain – Pathway has scientifically and formally mapped how intelligence emerges in the brain, enabling it to create BDH, an artificial reasoning system with a brain-like execution model.
BDH forms a modular structure similar to a network of neurons in the brain. It emerges spontaneously during training and resembles the behaviour of the neocortex, the outer layer of the brain present only in mammals, which is responsible for higher-order cognitive functions such as perception, memory, learning and decision-making.
In the BDH architecture, data inputs steer a population of artificial neurons, which build knowledge and draw inferences based on their interactions. This new approach enables a scale-free model, which can reason for long periods of time and behaves predictably even as new and unforeseen information is dynamically added to the model.
“To survive in a complex environment, humans need reasoning – based on both attention and language. That is why we focused on mapping natural human reasoning when solving the issue of generalization over time; trying to find a path from transformers to brain function,” said Zuzanna Stamirowska, CEO and co-founder of Pathway.
“We discovered that to achieve generalization over time we needed a completely new architecture. That’s why BDH is not an incremental improvement on transformer-based architecture, but a paradigm shift,” said Adrian Kosowski, Co-founder and Chief Scientific Officer at Pathway.
The research details how Pathway’s BDH architecture supports a number of benefits over transformer-based architectures. These include, but are not limited to:
- Generalization over time. The increased length of the chain-of-thought supports generalization over time, overcoming a significant barrier to autonomous intelligence.
- Predictability. Unlike today’s ‘black box’ systems, BDH ensures a provable risk level. The scale-free nature of BDH means that it will continue to reason in the way you expect over a long period.
- Safety. Shows how to overcome Bostrom’s famous ‘Paperclip Factory’ thought experiment (the risk of a superintelligent AI running on a ‘harmless goal’ malfunctioning and causing harm when running autonomously for a long time).
- Composability. Multiple BDH-based systems can be ‘glued’ together, resulting in emergent capabilities, similar to how a bilingual child develops fluency across two languages.
- Scarce data. The increased length of the chain-of-thought compared to transformer-based architectures means that BDH can reason for a longer time and draw better inferences on smaller datasets.
Competitive performance: BDH not only performs competitively on general-purpose hardware but it also has the potential for faster inference on specialised AI processors. This, in turn, has the potential to reduce the cost of reasoning and AI in enterprise deployments and most dramatically to generate tokens in a single long-running model with significantly lower latency.
“Generalization over time is the foundation for safe and autonomous reasoning,” said Stamirowska. “With BDH, we now have a scalable model that can sustain long-horizon reasoning, expanding the AI market in enterprise.”

