Can a 27B-Parameter AI Beat Frontier Models at Science? Inherent’s Faraday Raises the Stakes

Artificial intelligence is moving from systems that answer questions toward systems that can investigate them. The latest example comes from Inherent, a London-based AI laboratory founded by former Google DeepMind researchers, which says its Faraday AI agent has outperformed substantially larger frontier models from Anthropic and OpenAI on the demanding task of independently reproducing scientific research.
The significance of the development extends beyond a benchmark comparison. Faraday is designed around a more ambitious idea: an AI system should not merely retrieve scientific knowledge or generate plausible explanations, but should be capable of deciding what to investigate, designing experiments, executing them, evaluating the results and learning from the outcome.
That distinction could become important in the evolution of AI for scientific discovery. If systems can increasingly perform meaningful research with less human supervision, the bottleneck in scientific progress could shift from generating hypotheses to deciding which questions deserve computational and experimental resources.
Why Scientific Replication Is an Important AI Test
Scientific replication may appear less glamorous than discovering an entirely new theory, but it represents a demanding test of whether an AI system can actually perform research.
A scientific paper contains more than a conclusion. Reproducing its findings can require interpreting methodology, understanding experimental assumptions, locating or generating appropriate data, selecting tools, writing or adapting code, troubleshooting unexpected outcomes and determining whether the resulting evidence genuinely supports the published result.
This makes replication fundamentally different from ordinary question answering.
A language model can explain an experiment without actually conducting it. An autonomous research agent must bridge the gap between understanding an instruction and producing evidence.
Inherent has positioned Faraday around this distinction. The company's longer-term objective is to create AI capable of contributing to scientific discovery, and replication provides an intermediate test of whether an agent can operate within the practical workflow of research.
For human scientists, replication can also serve as foundational training. Researchers learn by studying existing work, reproducing results and gradually developing the judgment required to decide which experiments are worthwhile. Inherent is effectively attempting to encode a similar progression into an AI system.
Faraday’s Smaller Model Challenges the Bigger-Is-Better Assumption
One of the most notable aspects of Inherent's reported result is the underlying model used by Faraday.
According to the company, Faraday operates on Qwen 3.6, a model with approximately 27 billion parameters. It was compared with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5, described as much larger frontier-scale systems.
The comparison is significant because the AI industry has spent years demonstrating the benefits of scaling model size, training compute and infrastructure. Larger systems can provide greater general capability, but size alone does not guarantee superior performance on every specialized task.
Faraday illustrates another possibility: the surrounding agent architecture and training methodology may matter as much as the underlying foundation model.
Dimension | Conventional frontier approach | Inherent’s Faraday approach |
Core model | Large general-purpose model | Relatively compact 27B-parameter model |
Primary objective | Broad capability | Scientific research replication |
Training emphasis | General intelligence and task performance | Research-oriented reinforcement learning |
Research capability | Reasoning and tool use | Experiment selection and execution |
Development philosophy | Build broad capabilities | Focus deeply on scientific workflow |
This does not establish that smaller models are universally superior. Instead, it suggests that specialized training and agent design can potentially close or even reverse performance gaps on carefully selected tasks.
That possibility has major implications for the economics of AI research. If smaller models can achieve high performance when paired with sophisticated training and orchestration, organizations may be able to deploy scientific agents at lower computational cost than would be possible with the largest models.
The Hardest Capability May Be “Research Taste”
Inherent has emphasized a concept it calls “research taste.” The phrase refers to the ability to recognize which experiments are worth conducting and how those experiments should be designed.
This is substantially more difficult than following instructions.
A conventional automated system can execute a predefined procedure. A research agent must deal with uncertainty. Several possible experiments may be technically valid, but only some may provide meaningful information. Some experiments may be expensive, redundant or unlikely to resolve the underlying question.
Research therefore requires prioritization.
Consider a simplified scientific workflow:
Identify an unresolved question.
Form competing hypotheses.
Determine what evidence could distinguish them.
Select an experiment with useful information value.
Execute the experiment.
Analyze unexpected results.
Revise the hypothesis.
Determine the next experiment.
The important capability is not simply completing each step. It is deciding how to move between them.
Reinforcement learning is central to Inherent's approach because it can reward desirable outcomes rather than requiring developers to explicitly encode every rule governing scientific reasoning. This approach potentially allows an agent to discover strategies that would be difficult to specify manually.
The long-term question is whether these learned behaviors generalize. An agent that performs well on replication tasks must eventually demonstrate that its judgment remains useful when confronted with unfamiliar scientific domains, ambiguous evidence and genuinely novel problems.
Why Reinforcement Learning Could Matter for AI Scientists
Reinforcement learning changes the relationship between an AI system and its objectives.
Traditional supervised training often teaches a model to reproduce examples. Reinforcement learning instead creates an environment in which actions can be evaluated according to outcomes.
For scientific research, this distinction is attractive because science itself is outcome-driven.
An experiment can succeed, fail, produce an unexpected observation or reveal that an original assumption was incorrect. A capable research agent needs to respond differently to each possibility.
This could eventually produce systems that learn strategies rather than merely memorize scientific procedures.
The approach also reflects an important principle in autonomous AI development: capabilities can emerge from the interaction between a foundation model, tools, feedback mechanisms and an environment.
Faraday's use of existing software reinforces this philosophy. Rather than attempting to build every component internally, Inherent reportedly has the agent use OpenAI's GPT-5.5 Codex for coding tasks. The strategy resembles the way human researchers work with established tools instead of reinventing software infrastructure for every project.
That decision allows Inherent to concentrate its engineering effort on what it considers the harder problem, developing an agent capable of meaningful scientific judgment.
From AI Assistant to AI Research Teammate
The distinction between an assistant and a teammate may become increasingly important as AI agents become more autonomous.
An assistant generally waits for instructions. A research teammate can identify something interesting independently and return with evidence.
That behavior requires initiative, but initiative without discipline can be dangerous. An autonomous scientific system must distinguish productive exploration from wasted computation and plausible hypotheses from unsupported conclusions.
The ideal system would therefore operate within a collaborative loop:
Human defines the broad research objective → AI proposes investigations → AI executes experiments → AI evaluates evidence → human reviews findings → AI continues based on feedback.
Such a model could increase scientific productivity without requiring humans to surrender final responsibility for consequential conclusions.
The most valuable AI researcher may not be the system that produces the most text. It may be the system that discovers an overlooked connection, runs a decisive experiment and returns with evidence that changes the direction of a research program.
What Inherent’s Strategy Says About the AI Industry
Inherent's approach also represents a broader change in startup strategy.
The company emerged from stealth with a $50 million seed round and has operated with a relatively small team. Its reported plan is to grow from roughly a dozen employees to approximately 20 to 25 by the end of 2026.
This is a very different organizational model from the enormous teams and infrastructure budgets associated with frontier AI laboratories.
The strategy is based on specialization. Instead of attempting to compete with major AI companies across every capability, Inherent is concentrating on scientific agents and related research.
That specialization could become increasingly attractive as foundation models become more widely available. When high-quality models can be accessed externally, startups may be able to build differentiated products through training methods, agent architectures, proprietary environments, workflows and domain expertise rather than by training foundation models from scratch.
The result could be a more fragmented AI ecosystem in which small teams compete with large laboratories by solving narrowly defined, high-value problems exceptionally well.
London’s AI Talent Ecosystem Faces Its Own Challenges
Inherent's story is also connected to the development of London as a major AI center.
The company operates from King's Cross, an area strongly associated with the growth of Google's DeepMind presence and the broader concentration of artificial intelligence talent in London.
However, talent mobility remains a strategic issue.
Inherent cofounder Edward Hughes has criticized the UK's garden leave practices, which can restrict employees from immediately joining or establishing competing businesses after leaving an employer. From a startup perspective, delays in accessing experienced researchers can matter enormously in an industry where technical teams are often the primary competitive asset.
The movement of researchers from established laboratories into startups can therefore influence where new AI companies emerge and how quickly they can build.
Inherent's founders represent this broader talent-transfer phenomenon, moving experience gained at a leading AI research organization into a new company with a different scientific agenda.
The Potential Impact on Scientific Discovery
If AI agents eventually become capable of conducting reliable research with limited supervision, their impact could extend far beyond technology companies.
Potential applications include:
Drug and materials discovery
Climate and energy research
Physics simulation
Biological experimentation
Mathematical research
Cybersecurity research
Financial modeling
Advanced engineering
Quantum computing research
The common factor is not the subject matter itself. It is the presence of problems where large numbers of hypotheses can be evaluated through computational or experimental workflows.
AI could accelerate these workflows by allowing researchers to run more investigations in parallel.
However, greater experimentation also creates a new challenge: verification. Scientific progress depends on reproducibility, transparent methodology and independent validation. An AI system capable of generating thousands of hypotheses could increase the volume of scientific output without necessarily increasing the amount of reliable knowledge.
The future of AI-driven science will therefore depend on both discovery and verification.
The Critical Questions Faraday Still Has to Answer
Inherent's reported result is promising, but the broader scientific significance of Faraday will depend on what comes next.
Several questions remain central to evaluating autonomous research agents:
Generalization: Can Faraday reproduce research across substantially different scientific fields?
Novel discovery: Can it generate findings that were not already present in its training environment?
Reliability: How consistently can it distinguish genuine discoveries from experimental artifacts?
Cost efficiency: Can specialized agents deliver meaningful research at economically sustainable computational costs?
Human collaboration: Can researchers understand, audit and effectively guide the agent's reasoning?
Reproducibility: Can independent scientists reproduce discoveries made by autonomous agents?
These questions matter more than a single model comparison. A benchmark win demonstrates capability, but a scientific revolution requires durable performance under conditions where the correct answer is genuinely unknown.
The Bigger Picture for Predictive AI
The emergence of systems such as Faraday fits into a larger transition from predictive AI toward agentic intelligence.
Predictive systems estimate what is likely to happen. Generative systems create content. Agentic systems add planning, tool use, iteration and action.
Scientific research represents one of the most demanding environments for this evolution because the system cannot simply optimize for a conversationally satisfying response. It must interact with evidence.
For organizations working across advanced AI, quantum computing, financial modeling, cybersecurity and other complex domains, this distinction is particularly important. Research agents could eventually become a connective layer between massive datasets, simulations, specialized models and human decision-makers.
The strategic opportunity is not simply to automate researchers. It is to increase the number of scientifically meaningful questions that a research team can investigate.
The Race Is Moving From Smarter Models to Better Scientific Agents
Inherent's Faraday represents an intriguing development in the race toward AI capable of conducting research. Its reported performance against larger systems suggests that model scale is only one part of the equation. Specialized reinforcement learning, agent architecture, tool integration and research-oriented objectives may be equally important.
The more consequential development would be Faraday's progression from reproducing established findings to generating reliable new knowledge.
That transition will require much more than benchmark performance. It will require scientific judgment, rigorous verification, transparent reasoning, effective human collaboration and the ability to operate successfully when there is no known answer.
For the AI industry, that creates a compelling new frontier. The competition may increasingly be measured not by which model can answer the hardest question, but by which AI system can determine what question should be asked next.
For researchers and technology leaders, including the expert team at 1950.ai and Dr. Shahid Masood, the rise of autonomous scientific agents highlights a fundamental shift in advanced AI. Intelligence is becoming increasingly connected to experimentation, feedback and action. If that trajectory continues, AI could evolve from a tool that helps scientists work faster into a research partner capable of expanding the boundaries of what science can investigate.
Further Reading / External References
Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research
Inherent AI Faraday Replication, DeepMind Alumni





Comments