Frontier model capability
Independent NIST testing finds a measurable U.S. lead across cyber, software engineering, science, reasoning, and mathematics rather than inferring quality from compute alone.
The United States retains the current AI edge because independently tested frontier models are more capable, advanced compute and cloud ecosystems are deeper, and eight leading providers are being connected to classified military networks. The edge is narrower than compute or investment headlines imply: PRC models are improving quickly and can be cheaper at similar capability, China's research and talent scale is enormous, and neither side has publicly demonstrated a broad, repeatable military decision-performance advantage under realistic attack and data constraints.
Model quality is measured directly and kept separate from enabling inputs. NIST's independent evaluation places the leading tested PRC model about eight months behind the U.S. frontier, while also finding it cheaper than a comparable U.S. reference on five of seven cost tests. Compute, research, access, operational adoption, mission outcome, and assurance are therefore adjudicated independently. Classified-network agreements count as access—not decision superiority—and announced AI features count only when operational effect is observed.
Independent NIST testing finds a measurable U.S. lead across cyber, software engineering, science, reasoning, and mathematics rather than inferring quality from compute alone.
DeepSeek V4 was cheaper than a similar-capability U.S. reference on five of seven tested workloads, making deployment economics a separate competitive dimension.
Public evidence does not show which side consistently improves decision time, accuracy, workload, or mission effect with AI in representative contested operations.
Pair access to competing models with mission datasets, operator trials, security testing, and post-deployment monitoring; pay for repeatable outcome gains, not tokens or demonstrations.
Drivers are adjudicated separately so parallel strengths, dependencies, bottlenecks, and contrary evidence remain visible. The overall call is not a mechanical average.
Independent CAISI testing places the strongest evaluated PRC model about eight months behind the U.S. frontier across five capability domains.
The lag is an estimate from a bounded benchmark suite, not a universal intelligence measure; PRC models perform near the U.S. frontier on several individual tasks.
The U.S. and accessible allies control the deeper frontier accelerator, cloud, software, and model-provider stack, while the PRC remains constrained in access to leading accelerators.
PRC firms improve hardware/software efficiency, stockpile components, pursue domestic alternatives, and use intermediaries; aggregate U.S. compute does not equal military access at the edge.
DeepSeek V4 demonstrates that a PRC model can deliver similar aggregate capability at lower end-to-end cost on most tested workloads, supporting faster diffusion despite a frontier-quality lag.
The cost relationship varied by benchmark and excluded two evaluations; developer pricing, hosting configuration, reliability, and security can change total mission cost.
The United States retains disproportionate high-impact science, patents, global collaboration, and private-sector translation, while China leads publication volume, doctorate output, and several patent counts.
Broad science indicators are imperfect AI proxies; China surpassed the United States in comparable R&D volume and is building a formidable talent and innovation base.
Agreements with eight frontier AI companies for IL6 and IL7 networks create a uniquely broad, competitive pathway to classified operational use.
Agreements and network availability are not user adoption, mission integration, uptime, or verified effect; PLA military-civil fusion may move commercial advances through different channels.
Both militaries are integrating AI into analysis, cyber, unmanned systems, and planning, but public evidence does not establish a comparative, mission-specific improvement in decisions or outcomes.
Classified adoption may be substantial; PLA doctrine and U.S. provider access show serious institutional commitment even where outcomes remain undisclosed.
The United States has a stronger public evaluation and measurement-science ecosystem, but benchmark validity, deployment monitoring, cyber resilience, edge performance, and human-machine failure remain immature.
NIST's own work emphasizes that common evaluation and monitoring practices are preliminary and fragmented; a strong standards ecosystem does not prove fielded systems are safe or robust.
There is no single compute pipeline. Model capability, efficiency, research and infrastructure, secure military access, mission outcome, and assurance can move independently. Each thread shows which evidence is a direct outcome and which is only an enabling input.
Consequence-selected mission or capability chains drive the call.
Each fact is connected to a rule, bucket, and comparative effect.
Missing or incomparable evidence stays visible instead of becoming zero.
Mission effectProvide higher-quality reasoning, coding, cyber, scientific, perception, and agentic assistance than the opposing ecosystem.
Analytical scopePublicly released models evaluated under controlled scaffolding and token budgets; classified models and undocumented deployments remain unknown.
Action priorityPreserve and measure the edge
U.S. frontier quality leads in the strongest available independent comparison, but PRC efficiency materially narrows the operational and diffusion advantage.
Competition across providers, open-weight models, inference optimization, and workload-specific routing can lower U.S. cost without abandoning higher-capability options.
No benchmark suite fully represents military judgment, multimodal sensing, classified data, deception, human teaming, or long-horizon operations.
CAISI's April 2026 evaluation estimated that DeepSeek V4, the strongest PRC model it had tested, lagged the leading U.S. frontier by about eight months across cyber, software engineering, natural sciences, abstract reasoning, and mathematics.
~8 months estimated frontier lagA precommitted, multi-domain independent evaluation with held-out tasks receives greater weight than self-reported benchmarks or input proxies.
Directly supports a U.S. lead in model quality.
The time-lag estimate depends on selected models, benchmarks, scaffolding, and a fitted trend.
In CAISI's table, the leading U.S. model scored 71 percent versus DeepSeek V4's estimated 32 percent on a cyber benchmark and 78 versus 44 percent on the held-out PortBench software task, while results were close on several science and mathematics tests.
71–32 / 78–44 percentage-point benchmark comparisonsLarge differences on held-out, mission-relevant task families establish a direct capability edge; near-parity elsewhere prevents universalizing it.
Shows where the model lead is consequential and where it is not.
Benchmarks are proxies and the cyber value for DeepSeek was imputed from a subset of samples.
DeepSeek V4 was less expensive than the selected similar-capability U.S. reference on five of seven comparable benchmarks; its relative cost ranged from 53 percent lower to 41 percent higher.
5 of 7 benchmarks with lower DeepSeek costWorkload-level cost per successful outcome is a distinct competitive measure and cannot be inferred from token price alone.
Reduces the practical size of the U.S. capability lead.
Two benchmarks were excluded and provider prices, hardware, and serving configurations can change.
Mission effectTrain, serve, improve, and diffuse capable AI systems while replacing constrained hardware, energy, data, and talent inputs.
Analytical scopeAdvanced accelerators, cloud, R&D, high-impact output, talent, public research access, and substitution—not raw compute or publication count alone.
Action priorityKeep the ecosystem open and resilient
The U.S. ecosystem retains frontier hardware, highly cited research, and business-led translation advantages, while China's scale in R&D, publications, doctorates, patents, and efficiency makes the contest much closer than a chip-control narrative suggests.
Allied semiconductor depth, NAIRR access, energy and data-center expansion, talent attraction, and research collaboration expand the accessible U.S. base.
Comparable frontier training compute, usable military data, researcher flows, energy constraints, and domestic chip yield are incomplete or proprietary.
The 2025 Defense Department report assesses that China's AI sector remains constrained by limited access to high-performance accelerators while pursuing efficiency, stockpiles, circumvention, and domestic alternatives.
A documented chokepoint supports a relative edge only when the competitor's active substitutions and observed model performance are also considered.
Supports a U.S. compute edge but not a complete AI lead.
Stockpiles, illicit access, domestic output, and actual training allocations are uncertain.
NSF's 2026 indicators estimate China produced 31 percent of global science and engineering articles in 2024 versus 12 percent for the United States, while U.S. work retained a disproportionate share of highly cited output and internationally coauthored articles.
31 / 12 percent of global S&E articlesVolume, citation impact, and collaboration are separate measures; high volume cannot be equated with frontier quality, and high impact cannot erase competitor scale.
Holds the talent and research driver to a narrow U.S. edge.
All-science publications are a broad AI ecosystem proxy and citation measures lag current work.
The NSF-led NAIRR reported supporting more than 600 research teams and 6,000 students across all U.S. states, Washington, D.C., and Puerto Rico through public-private compute, data, models, and expertise.
600 / 6000 research teams / studentsObserved access to advanced resources and training is a current ecosystem enabler, not direct model or military performance.
Broadens the U.S. innovation base beyond frontier firms.
Participation does not establish research quality, military relevance, or comparison with PRC programs.
Mission effectMove frontier and specialized AI into classified workflows with data, permissions, tools, and accountable human use.
Analytical scopeAvailable providers, network levels, data connections, user access, latency, training, acquisition, and sustained use—not contract count.
Action priorityConvert access into adoption
The United States has created a broad classified-access pathway across eight leading firms, but public evidence does not yet show user scale, mission availability, or operational effect.
Provider competition, portable evaluations, shared security services, and mission-level procurement can avoid lock-in and accelerate adoption.
Active users, approved use cases, model versions, uptime, data access, cost, and field effects on IL6/IL7 networks are not public.
In May 2026 the Department announced agreements with eight leading AI companies to make advanced capabilities available on Impact Level 6 and 7 classified networks.
8 frontier AI companiesMultiple providers on mission networks create a current access and competition advantage, while adoption and effect remain separate thresholds.
Supports a U.S. edge in secure deployment options.
The release does not report operational availability, user counts, model versions, cost, or mission outcome.
The 2025 PRC report assesses that military-civil fusion gives the PLA continuing access to commercial and academic AI advances and records AI-related work in unmanned systems, ISR analysis, decision assistance, cyber operations, and information campaigns.
A structured diffusion mechanism and observed applications establish serious adoption activity, but not comparative mission effect.
Prevents the U.S. access advantage from becoming a lead.
Specific fielded systems, users, performance, and security are not disclosed.
No public source establishes the share of U.S. or PLA operational units that routinely use AI for priority decisions, the tasks delegated, or the resulting performance.
Network availability, contracts, demonstrations, and doctrine cannot substitute for sustained use by trained operators in real workflows.
Caps confidence in the U.S. secure-access edge.
Mission effectReduce decision time or workload and improve accuracy, allocation, targeting, logistics, cyber response, or mission outcome without unacceptable error.
Analytical scopeContext-specific human-plus-AI performance against a baseline, including latency, confidence, override, failure, and downstream effect.
Action priorityMake the central claim measurable
Public evidence supports strong capabilities and active adoption pathways, but neither side demonstrates a broad current advantage in military decision outcomes.
Instrumented operator trials, shadow mode, controlled rollout, and mission-level A/B comparisons can produce direct evidence without exposing classified tactics.
Relevant baselines, tasks, error costs, decision authority, adversary adaptation, and operational outcomes are classified or not standardized.
The PRC report describes Multi-Domain Precision Warfare as using big data and AI to aggregate information, identify weak points, and support rapid operational decisions, while stating that PLA intelligentized-warfare theory and concepts remain under development and experimentation.
A coherent doctrine and recurring experimentation establish intent and learning, but not demonstrated decision advantage.
Shows the PRC is competing directly in decision systems.
The assessment does not publish unit adoption, decision metrics, or combat results.
The classified-network announcement states intended benefits in data synthesis, situational understanding, and warfighter decision support, but reports no measured baseline or achieved effect.
A stated benefit or plausible use case remains an objective until a controlled comparison shows time, quality, workload, or mission improvement.
Prevents model access from being scored as decision superiority.
Operational results may be classified or collected after the announcement.
No public dataset compares U.S./allied and PLA human-plus-AI decision time, accuracy, workload, error severity, and mission effect on representative national-security tasks.
When the defining outcome lacks comparable observations, it remains unknown rather than inherited from model, compute, or publication leadership.
Keeps operational decision performance contested.
Mission effectMaintain useful, secure, calibrated AI behavior as models, data, users, networks, and threats change.
Analytical scopeEvaluation validity, cyber and misuse risk, adversarial inputs, drift, disconnected operation, human override, monitoring, and recovery.
Action priorityTurn evaluation into an operational control loop
The U.S. measurement ecosystem is a competitive asset, but NIST's own findings show benchmark and post-deployment monitoring practice is still immature. There is no demonstrated comparative field-robustness lead.
Independent held-out evaluation, provider diversity, logging, red teaming, rollback, edge testing, and continuous mission monitoring reduce dependence on one benchmark or model.
Classified incident rates, adversarial success, drift, edge performance, operator misuse, and mission recovery are not publicly comparable.
NIST AI 800-2 describes current language-model and agent benchmark practices as preliminary and organizes them around defining the measurement target, implementing the evaluation, and analyzing/reporting results.
A transparent evaluation process improves decision quality, but preliminary voluntary practice is an enabler rather than field assurance.
Provides a modest U.S. institutional advantage.
The document is an initial public draft and automated benchmarks cover only part of assurance.
CAISI found DeepSeek V4 appeared roughly at frontier parity on developer-reported benchmarks but lagged on CAISI's precommitted suite, including held-out reasoning, software, and cyber tasks.
A material difference between developer-selected and independent held-out results requires procurement decisions to use independent, mission-specific evaluation.
Validates independent testing as a competitive control, not paperwork.
NIST AI 800-4 concludes that monitoring is crucial for reliability, unexpected behavior, and real-world consequences, while common methods, terminology, guidance, and information sharing remain nascent and fragmented.
6 monitoring categoriesIf validated monitoring practice is immature, pre-deployment benchmark success cannot be treated as sustained operational reliability.
Keeps contested robustness from becoming a U.S. edge.
The report identifies challenges and categories rather than grading specific military deployments.
Lead time is an implementation attribute—not a forecasted future rating. Every action has an owner, prerequisite, and observable completion test.
Frontier rankings and provider claims do not show which model, scaffold, data, and operator combination improves a specific military decision.
Maintain held-out mission suites spanning planning, cyber, intelligence, logistics, autonomy, and contested edge use; evaluate multiple U.S. and adversary models with controlled tools, token budgets, latency, cost, uncertainty, and human baselines.
PrerequisitesProtected datasets, independent evaluators, reproducible harnesses, provider access, and releasable summary methods.
Verify successEvery priority deployment has current independent results against a human and non-AI baseline, including cost, latency, failure, and adversarial conditions.
Eight-provider availability can remain shelfware if mission data, tools, permissions, training, and workflow ownership are missing.
Select consequential workflows; integrate data and tools; deploy in shadow mode; measure decision time, quality, workload, override, and downstream effect; scale only where a named owner accepts the result.
PrerequisitesIL6/IL7 access, mission-data contracts, workflow owners, operator training, and evaluation/monitoring pipelines.
Verify successAdopted systems deliver statistically and operationally meaningful improvement against baseline without unacceptable error or hidden manual burden.
Model leadership depends on compute, energy, chips, cloud, talent, data, tools, and high-impact research—not one export-control chokepoint.
Expand power and data-center delivery, allied semiconductor and cloud capacity, research-compute access, talent attraction, secure full-stack export packages, and rapid substitution plans while testing whether controls produce the intended capability effect.
PrerequisitesInfrastructure permits, energy supply, security agreements, research funding, immigration/talent policy, and outcome-based control review.
Verify successAccessible allied compute, model, and research capacity grows while critical single-source exposure and time to add power or replace constrained components decline.
Linked threadsAI-T2
AI behavior changes with model updates, data, prompts, users, adversary inputs, and infrastructure; pre-deployment approval decays quickly.
Require mission-specific telemetry, drift and incident detection, adversarial testing, user feedback, human override, safe degradation, model/version provenance, rollback, and decommissioning for every operational AI service.
PrerequisitesLogging standards, evaluation triggers, protected incident sharing, model inventory, and operational rollback authority.
Verify successDeployments detect representative drift and attack, revert safely, and restore approved performance within mission thresholds during exercises and live incidents.
Linked threadsAI-T5
The finding comes first. Open this section to inspect composition rules, evidence limitations, and every source record.
The assessment evaluates AI capability and decision advantage usable now using public evidence through August 5, 2026. Model quality, enabling inputs, access, adoption, mission performance, and assurance are separate dimensions.
Threads are included when they represent an independently varying source of advantage or failure: model capability/efficiency, ecosystem capacity, secure access, mission outcome, and robustness.
Direct independent performance and observed operational outcome outrank compute, publications, contracts, or announced use cases. The overall score is adjudicated and not a weighted input index.
Held-out government evaluation, official infrastructure and adoption records, harmonized science indicators, and measurement-science findings are separated from developer claims and strategic intent.
Classified models, mission datasets, user adoption, decision outcomes, incidents, and PLA performance remain unknown. Missing outcomes are not inferred from inputs on either side.
Primary use: Independent U.S./PRC frontier capability, task-level benchmark, cost, and evaluation-selection evidence.
Known limitation: Bounded public-model comparison whose results depend on benchmark suite, scaffolding, and serving choices.
Open public source ↗Primary use: Earlier cross-model capability, security, censorship, adoption, and cost evidence.
Known limitation: Superseded in part by newer model releases; remains useful for evaluation continuity and security dimensions.
Open public source ↗Primary use: Eight-provider IL6/IL7 access, intended use, and classified-network deployment pathway.
Known limitation: Official announcement does not report adoption, uptime, cost, model versions, or achieved mission effects.
Open public source ↗Primary use: PRC military AI applications, model progress, accelerator constraints, military-civil fusion, doctrine, and experimentation.
Known limitation: Unclassified threat assessment; PLA deployment, performance, and underlying methods are partly undisclosed.
Open public source ↗Primary use: Comparable R&D, publication, citation, patent, collaboration, and science-workforce indicators.
Known limitation: Broad and often lagged science indicators are not direct measures of frontier model or military performance.
Open public source ↗Primary use: Research-team, student, geographic, compute, data, and public-private access evidence.
Known limitation: Participation and resource access do not establish research quality, model leadership, or military relevance.
Open public source ↗Primary use: Evaluation objectives, implementation, analysis, reporting, and preliminary-practice limits.
Known limitation: Initial public draft focused on automated benchmarks, not complete operational assurance.
Open public source ↗Primary use: Post-deployment monitoring importance, categories, gaps, barriers, and immature practice.
Known limitation: Cross-sector research identifies challenges but does not evaluate named military deployments.
Open public source ↗Primary use: Governance, data, performance, and monitoring accountability framework for operational adoption.
Known limitation: Federal accountability framework predates the latest frontier models and is not a comparative U.S./PRC performance assessment.
Open public source ↗The public JSON contains the executive judgment, drivers, decisive threads, observations, actions, limitations, and source records.
This applied layer substantiates the current call without projecting future ratings or silently changing the core ten-area dataset. Compare the core capability record →