What Will Be the Next Key Battleground for Large Models After AI Coding?

08/07 2026 514

Over the past 18 months, the realm of AI programming has witnessed a frenetic transformation. The utilization rate of large models in coding has soared from less than 5% to over 88%. On authoritative platforms like MirrorCode, leading models consistently top the charts with an absolute success rate of 64%, while the scores of newcomers lag far behind, at only a third of theirs.

The cycle of code generation, testing, and refinement is accelerating rapidly. According to calculations by the AI investment research platform Funda, the combined annualized revenue of OpenAI and Anthropic is projected to reach approximately $120 billion within the next 12 months, with AI Coding serving as a pivotal growth driver.

While the coding "track" may appear to be ablaze with activity, a subtle shift is underway—numerous benchmark tests have been "maxed out" by all mainstream models. Leaderboards are losing their discriminatory power, and the scope for incremental growth is noticeably diminishing.

Meanwhile, another frontier remains largely untapped. In a joint evaluation of scientific research processes conducted by the Shanghai AI Laboratory and 100 scientists, the scores of the most advanced large models did not exceed 50 points. Among the 90 top-tier real-world scientific research tasks covered by NatureBench, only 17.8% of the optimal intelligent agents surpassed the results of original papers. Even more striking is another statistic: over the past 18 months, the completion rate of the best large models in end-to-end scientific discoveries has remained stagnant at 3%, with virtually no change.

On one hand, there is a stark contrast between the 88% utilization rate in coding and the mere 3% in scientific discoveries; on the other, there is a dichotomy between the "maxed out" performance in coding and the "almost unchanged" state in scientific research. Zhou Bowen, Chief Scientist at the Shanghai Artificial Intelligence Laboratory (Pujiang Laboratory), succinctly summarized the situation: "Research will be the next Coding." The focal point of competition among large models is shifting from coding to scientific research.

01. Why Did Coding Take the Lead?

To comprehend why this shift is occurring, we must first understand why coding emerged as the earliest field where large models delivered tangible value. The answer lies in a crucial factor: feedback.

Coding inherently possesses a comprehensive set of "validators"—compilation, testing, and runtime results. Each time a model generates a piece of code, it can swiftly ascertain its correctness. Compilation failed? Revise it. Test failed? Revise it again. A closed loop of "hypothesis-validation-correction" can be completed in mere seconds. This ability to rapidly falsify hypotheses provides large models with nearly flawless training signals in the programming domain.

Scientific research, however, presents an entirely different landscape. Zhou Bowen identified three structural impediments.

Firstly, large models can only passively observe the world, learning correlations rather than causality. They can summarize patterns from vast datasets but cannot grasp the underlying "why."

Secondly, the more challenging the problem, the scarcer the feedback. While coding can validate correctness in seconds, a scientific hypothesis may require months or even years of experimental testing, with failure being the norm.

Thirdly, the scientific research outcomes published by humans are predominantly successful data, devoid of failed attempts. Models are thus confined to outdated knowledge distributions, merely imitating without genuinely innovating. Evaluations like NatureBench further underscore this: even when provided with all references and data, the best models fail to replicate the results of a Nature-level paper.

02. Why Is Scientific Research the Next Frontier?

So why is scientific research, rather than other fields, poised to become the next battleground? The answer lies in the fact that scientific research scenarios inherently possess the foundational conditions to succeed coding.

From a commercial standpoint, scientific research represents a vastly larger market than programming. According to the European Commission, the top 2,000 companies globally in terms of R&D investment collectively invested approximately €1.446 trillion in 2024. Industries such as life sciences, new materials, chemicals, and energy allocate enormous sums to R&D annually, yet a significant portion of resources is still squandered on inefficient trial-and-error and repetitive experiments. If AI can enhance scientific research efficiency by even a marginal percentage, the value created would be astronomical.

From a technological perspective, scientific research is retracing the trajectory of coding. The "validator" for coding is the compiler, while the "validator" for scientific research is becoming real-world experiments. The "Design-Manufacture-Test-Analyze" (DMTA) closed loop is evolving into the "compiler" of the scientific realm—subjecting every model hypothesis to rigorous experimental validation and feeding the results back into the next round of predictions. When AI can enter laboratories, operate robotic arms, and complete autonomous scientific research in a "dry-wet closed loop," it will transcend being merely a brain proposing hypotheses in the digital world to becoming a complete research entity capable of validating hypotheses in the physical world.

03. Signals Are Already Emerging

This shift is not a distant prospect but is already underway.

In early August, OpenAI revealed that its test models had solved or contributed to solving 10 long-standing unsolved mathematical problems spanning high-dimensional geometry, coding theory, group theory, and other fields. DeepMind's AlphaProof Nexus successfully tackled 9 long-standing unsolved Erdős open problems, the earliest of which had baffled the academic community for 56 years.

The common thread among these breakthroughs is that AI is no longer just "calculating faster" but is beginning to "think further"—proposing hypotheses, exploring diverse paths, analyzing failures, summarizing conclusions, and continuously adjusting its direction accordingly.

In broader scientific research scenarios, infrastructure is being rapidly developed. The PanShi·Scientific Foundational Large Model 2.0, launched by the Chinese Academy of Sciences, aggregates 8 million high-quality scientific reasoning data points covering over 200 scientific research tasks.

The "ShuSheng·DuanYan" platform by the Shanghai AI Laboratory has integrated the complete scientific research process from hypothesis formulation to experimental validation. Anthropic released Claude Science, a scientific research agent workbench, while ByteDance launched the Seed STEM Scientist Program. XtalPi (02228.HK) unveiled an open intelligent R&D platform for scientific research scenarios. Leading companies are shifting from "competing on model parameters" to "securing a foothold in the scientific research ecosystem." As multiple industry institutions have predicted, future core opportunities in the AI industry will revolve around large models, intelligent agents, physical AI, and AI4S (AI for Science).

04. What Matters More Than Winning or Losing?

Of course, the leap from coding to scientific research is far from a smooth transition. Existing evaluations have yet to truly test AI's ability to "discover new methods." Foundational innovations that have genuinely propelled progress in machine learning—such as convolutions, residual connections, and Attention—were not achieved by scoring high on fixed leaderboards.

Evaluations like MLS-Bench, jointly proposed by research teams including the University of California, Berkeley, demonstrate that even when cutting-edge models are given complete implementations of the best human methods and allowed multiple rounds of experimentation, the methods they propose still do not outperform humans overall. Models excel primarily at fine-tuning and recombining existing methods, with truly novel methodological mechanisms remaining scarce.

But this precisely underscores the deeper significance of this competition. Zhou Bowen proposed a thought-provoking distinction: the ultimate goal of AI for Science should not be a "revolution of tools"—creating better tools—but rather a "revolutionary tool" capable of discovering new paradigms and helping humanity break through cognitive limitations. He envisions that if today's AI were placed in 1905, after studying Riemannian geometry and special relativity, it might independently derive general relativity.

This is the true essence of scientific research as the "next Coding"—it is not about repeating the known but about exploring the unknown.

Coding capabilities have taught AI "how to do things," while scientific research capabilities aim to teach AI "what to do" and "why to do it." As large models transition from compilers to laboratories, from instantaneous feedback to prolonged trial-and-error, from replicating papers to proposing hypotheses, the challenges they face transcend engineering efficiency and delve into the deep waters of scientific cognition.

The endpoint of this competition is not who can write better code but who can first reach that "scientific meta-cognitive moment"—where AI knows what it does not know.

- End -

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.