← Capability assessment Applied current-state analysis

AI & decision

The United States retains the current AI edge because independently tested frontier models are more capable, advanced compute and cloud ecosystems are deeper, and eight leading providers are being connected to classified military networks. The edge is narrower than compute or investment headlines imply: PRC models are improving quickly and can be cheaper at similar capability, China's research and talent scale is enormous, and neither side has publicly demonstrated a broad, repeatable military decision-performance advantage under realistic attack and data constraints.

Evidence cutoff · August 5, 2026 5 decisive threads 15 applied observations 4 ranked actions
Current comparative call

U.S. edge

Moderate confidence
PRC leadContestedU.S. / credible-allied lead
Why this call

Model quality is measured directly and kept separate from enabling inputs. NIST's independent evaluation places the leading tested PRC model about eight months behind the U.S. frontier, while also finding it cheaper than a comparable U.S. reference on five of seven cost tests. Compute, research, access, operational adoption, mission outcome, and assurance are therefore adjudicated independently. Classified-network agreements count as access—not decision superiority—and announced AI features count only when operational effect is observed.

Strongest U.S. advantage

Frontier model capability

Independent NIST testing finds a measurable U.S. lead across cyber, software engineering, science, reasoning, and mathematics rather than inferring quality from compute alone.

Most important PRC offset

Efficiency and rapid diffusion

DeepSeek V4 was cheaper than a similar-capability U.S. reference on five of seven tested workloads, making deployment economics a separate competitive dimension.

Largest analytical gap

Military decision performance

Public evidence does not show which side consistently improves decision time, accuracy, workload, or mission effect with AI in representative contested operations.

Decision

Buy measured mission improvement

Pair access to competing models with mission datasets, operator trials, security testing, and post-deployment monitoring; pay for repeatable outcome gains, not tokens or demonstrations.

Current advantage profile

Where the edge actually sits.

Drivers are adjudicated separately so parallel strengths, dependencies, bottlenecks, and contrary evidence remain visible. The overall call is not a mechanical average.

AI-D1 · High confidence

Model quality & demonstrated capability

U.S. lead
PRC advantageU.S. / credible-allied advantage

Independent CAISI testing places the strongest evaluated PRC model about eight months behind the U.S. frontier across five capability domains.

Contrary evidence and caveat

The lag is an estimate from a bounded benchmark suite, not a universal intelligence measure; PRC models perform near the U.S. frontier on several individual tasks.

AI-D2 · Moderate confidence

Compute, chips & cloud infrastructure

U.S. edge
PRC advantageU.S. / credible-allied advantage

The U.S. and accessible allies control the deeper frontier accelerator, cloud, software, and model-provider stack, while the PRC remains constrained in access to leading accelerators.

Contrary evidence and caveat

PRC firms improve hardware/software efficiency, stockpile components, pursue domestic alternatives, and use intermediaries; aggregate U.S. compute does not equal military access at the edge.

AI-D3 · Moderate confidence

Cost, efficiency & diffusion

PRC edge
PRC advantageU.S. / credible-allied advantage

DeepSeek V4 demonstrates that a PRC model can deliver similar aggregate capability at lower end-to-end cost on most tested workloads, supporting faster diffusion despite a frontier-quality lag.

Contrary evidence and caveat

The cost relationship varied by benchmark and excluded two evaluations; developer pricing, hosting configuration, reliability, and security can change total mission cost.

AI-D4 · Low–moderate confidence

Talent, research & data ecosystem

U.S. edge
PRC advantageU.S. / credible-allied advantage

The United States retains disproportionate high-impact science, patents, global collaboration, and private-sector translation, while China leads publication volume, doctorate output, and several patent counts.

Contrary evidence and caveat

Broad science indicators are imperfect AI proxies; China surpassed the United States in comparable R&D volume and is building a formidable talent and innovation base.

AI-D5 · Moderate confidence

Secure model access & deployment pathways

U.S. edge
PRC advantageU.S. / credible-allied advantage

Agreements with eight frontier AI companies for IL6 and IL7 networks create a uniquely broad, competitive pathway to classified operational use.

Contrary evidence and caveat

Agreements and network availability are not user adoption, mission integration, uptime, or verified effect; PLA military-civil fusion may move commercial advances through different channels.

AI-D6 · Low confidence

Operational integration & decision performance

Contested
PRC advantageU.S. / credible-allied advantage

Both militaries are integrating AI into analysis, cyber, unmanned systems, and planning, but public evidence does not establish a comparative, mission-specific improvement in decisions or outcomes.

Contrary evidence and caveat

Classified adoption may be substantial; PLA doctrine and U.S. provider access show serious institutional commitment even where outcomes remain undisclosed.

AI-D7 · Low–moderate confidence

Evaluation, assurance & contested robustness

Contested
PRC advantageU.S. / credible-allied advantage

The United States has a stronger public evaluation and measurement-science ecosystem, but benchmark validity, deployment monitoring, cyber resilience, edge performance, and human-machine failure remain immature.

Contrary evidence and caveat

NIST's own work emphasizes that common evaluation and monitoring practices are preliminary and fragmented; a strong standards ecosystem does not prove fielded systems are safe or robust.

Applied causal analysis

Five parallel AI threads determine usable decision advantage.

There is no single compute pipeline. Model capability, efficiency, research and infrastructure, secure military access, mission outcome, and assurance can move independently. Each thread shows which evidence is a direct outcome and which is only an enabling input.

Decisive threads5

Consequence-selected mission or capability chains drive the call.

Applied observations15

Each fact is connected to a rule, bucket, and comparative effect.

Unknown / unscored2

Missing or incomparable evidence stays visible instead of becoming zero.

01 · AI-T1
Frontier model capability and efficiencyIndependent task performance at mission-relevant cost
U.S. edgeHigh confidence

Mission effectProvide higher-quality reasoning, coding, cyber, scientific, perception, and agentic assistance than the opposing ecosystem.

Analytical scopePublicly released models evaluated under controlled scaffolding and token budgets; classified models and undocumented deployments remain unknown.

Action priorityPreserve and measure the edge

  1. 01Capability
  2. 02Cost efficiency
  3. 03Mission fit
  4. 04Operational value
Judgment

U.S. frontier quality leads in the strongest available independent comparison, but PRC efficiency materially narrows the operational and diffusion advantage.

What offsets it

Competition across providers, open-weight models, inference optimization, and workload-specific routing can lower U.S. cost without abandoning higher-capability options.

What remains unknown

No benchmark suite fully represents military judgment, multimodal sensing, classified data, deception, human teaming, or long-horizon operations.

Observation → applied rule → condition → driver effect3 observations
AI-O01 · Aggregate capability

CAISI's April 2026 evaluation estimated that DeepSeek V4, the strongest PRC model it had tested, lagged the leading U.S. frontier by about eight months across cyber, software engineering, natural sciences, abstract reasoning, and mathematics.

~8 months estimated frontier lag
Independent frontier capability

A precommitted, multi-domain independent evaluation with held-out tasks receives greater weight than self-reported benchmarks or input proxies.

U.S. frontier lead

Directly supports a U.S. lead in model quality.

April 2026 evaluation · High confidence

The time-lag estimate depends on selected models, benchmarks, scaffolding, and a fitted trend.

AI-O02 · Task-level performance

In CAISI's table, the leading U.S. model scored 71 percent versus DeepSeek V4's estimated 32 percent on a cyber benchmark and 78 versus 44 percent on the held-out PortBench software task, while results were close on several science and mathematics tests.

71–32 / 78–44 percentage-point benchmark comparisons
Domain-specific demonstrated capability

Large differences on held-out, mission-relevant task families establish a direct capability edge; near-parity elsewhere prevents universalizing it.

U.S. edge with domain variation

Shows where the model lead is consequential and where it is not.

April 2026 evaluation · High confidence

Benchmarks are proxies and the cyber value for DeepSeek was imputed from a subset of samples.

AI-O03 · Cost efficiency

DeepSeek V4 was less expensive than the selected similar-capability U.S. reference on five of seven comparable benchmarks; its relative cost ranged from 53 percent lower to 41 percent higher.

5 of 7 benchmarks with lower DeepSeek cost
End-to-end cost for correctly solved tasks

Workload-level cost per successful outcome is a distinct competitive measure and cannot be inferred from token price alone.

PRC efficiency offset

Reduces the practical size of the U.S. capability lead.

April 2026 evaluation · High confidence

Two benchmarks were excluded and provider prices, hardware, and serving configurations can change.

02 · AI-T2
Compute, talent and research ecosystemSustain frontier development and broad adoption
U.S. edgeModerate confidence

Mission effectTrain, serve, improve, and diffuse capable AI systems while replacing constrained hardware, energy, data, and talent inputs.

Analytical scopeAdvanced accelerators, cloud, R&D, high-impact output, talent, public research access, and substitution—not raw compute or publication count alone.

Action priorityKeep the ecosystem open and resilient

  1. 01Compute & energy
  2. 02Talent & research
  3. 03Data & tools
  4. 04Frontier iteration
Judgment

The U.S. ecosystem retains frontier hardware, highly cited research, and business-led translation advantages, while China's scale in R&D, publications, doctorates, patents, and efficiency makes the contest much closer than a chip-control narrative suggests.

What offsets it

Allied semiconductor depth, NAIRR access, energy and data-center expansion, talent attraction, and research collaboration expand the accessible U.S. base.

What remains unknown

Comparable frontier training compute, usable military data, researcher flows, energy constraints, and domestic chip yield are incomplete or proprietary.

Observation → applied rule → condition → driver effect3 observations
AI-O04 · Advanced accelerators

The 2025 Defense Department report assesses that China's AI sector remains constrained by limited access to high-performance accelerators while pursuing efficiency, stockpiles, circumvention, and domestic alternatives.

Frontier compute access

A documented chokepoint supports a relative edge only when the competitor's active substitutions and observed model performance are also considered.

PRC compute constraint

Supports a U.S. compute edge but not a complete AI lead.

2024 activity reported in 2025 · Moderate–high confidence

Stockpiles, illicit access, domestic output, and actual training allocations are uncertain.

AI-O05 · Research scale and impact

NSF's 2026 indicators estimate China produced 31 percent of global science and engineering articles in 2024 versus 12 percent for the United States, while U.S. work retained a disproportionate share of highly cited output and internationally coauthored articles.

31 / 12 percent of global S&E articles
Research scale and impact

Volume, citation impact, and collaboration are separate measures; high volume cannot be equated with frontier quality, and high impact cannot erase competitor scale.

Split research advantage

Holds the talent and research driver to a narrow U.S. edge.

2024 output reported in 2026 · High confidence

All-science publications are a broad AI ecosystem proxy and citation measures lag current work.

AI-O06 · Research access

The NSF-led NAIRR reported supporting more than 600 research teams and 6,000 students across all U.S. states, Washington, D.C., and Puerto Rico through public-private compute, data, models, and expertise.

600 / 6000 research teams / students
Distributed research-resource access

Observed access to advanced resources and training is a current ecosystem enabler, not direct model or military performance.

U.S. diffusion strength

Broadens the U.S. innovation base beyond frontier firms.

March 2026 · High confidence

Participation does not establish research quality, military relevance, or comparison with PRC programs.

03 · AI-T3
Secure government access and adoptionOperators can use competing models on mission networks
U.S. edgeModerate confidence

Mission effectMove frontier and specialized AI into classified workflows with data, permissions, tools, and accountable human use.

Analytical scopeAvailable providers, network levels, data connections, user access, latency, training, acquisition, and sustained use—not contract count.

Action priorityConvert access into adoption

  1. 01Provider access
  2. 02Mission data
  3. 03User workflow
  4. 04Sustained adoption
Judgment

The United States has created a broad classified-access pathway across eight leading firms, but public evidence does not yet show user scale, mission availability, or operational effect.

What offsets it

Provider competition, portable evaluations, shared security services, and mission-level procurement can avoid lock-in and accelerate adoption.

What remains unknown

Active users, approved use cases, model versions, uptime, data access, cost, and field effects on IL6/IL7 networks are not public.

Observation → applied rule → condition → driver effect3 observations
AI-O07 · Classified access

In May 2026 the Department announced agreements with eight leading AI companies to make advanced capabilities available on Impact Level 6 and 7 classified networks.

8 frontier AI companies
Competitive classified model access

Multiple providers on mission networks create a current access and competition advantage, while adoption and effect remain separate thresholds.

Broad U.S. access pathway

Supports a U.S. edge in secure deployment options.

May 2026 · High confidence

The release does not report operational availability, user counts, model versions, cost, or mission outcome.

AI-O08 · PLA commercial integration

The 2025 PRC report assesses that military-civil fusion gives the PLA continuing access to commercial and academic AI advances and records AI-related work in unmanned systems, ISR analysis, decision assistance, cyber operations, and information campaigns.

Commercial-to-military diffusion

A structured diffusion mechanism and observed applications establish serious adoption activity, but not comparative mission effect.

PRC adoption pathway

Prevents the U.S. access advantage from becoming a lead.

2024 activity reported in 2025 · Moderate confidence

Specific fielded systems, users, performance, and security are not disclosed.

AI-O09 · Sustained operational adoption

No public source establishes the share of U.S. or PLA operational units that routinely use AI for priority decisions, the tasks delegated, or the resulting performance.

Adoption at operational scale

Network availability, contracts, demonstrations, and doctrine cannot substitute for sustained use by trained operators in real workflows.

Unknown comparison

Caps confidence in the U.S. secure-access edge.

August 2026 · High confidence

Classified adoption may be substantial on both sides.

04 · AI-T4
Mission decision performanceObserved improvement in a defined operational decision
ContestedLow confidence

Mission effectReduce decision time or workload and improve accuracy, allocation, targeting, logistics, cyber response, or mission outcome without unacceptable error.

Analytical scopeContext-specific human-plus-AI performance against a baseline, including latency, confidence, override, failure, and downstream effect.

Action priorityMake the central claim measurable

  1. 01Mission task
  2. 02Human–AI workflow
  3. 03Decision quality
  4. 04Mission outcome
Judgment

Public evidence supports strong capabilities and active adoption pathways, but neither side demonstrates a broad current advantage in military decision outcomes.

What offsets it

Instrumented operator trials, shadow mode, controlled rollout, and mission-level A/B comparisons can produce direct evidence without exposing classified tactics.

What remains unknown

Relevant baselines, tasks, error costs, decision authority, adversary adaptation, and operational outcomes are classified or not standardized.

Observation → applied rule → condition → driver effect3 observations
AI-O10 · PLA decision concept

The PRC report describes Multi-Domain Precision Warfare as using big data and AI to aggregate information, identify weak points, and support rapid operational decisions, while stating that PLA intelligentized-warfare theory and concepts remain under development and experimentation.

Operational concept maturity

A coherent doctrine and recurring experimentation establish intent and learning, but not demonstrated decision advantage.

Serious development / effect unproven

Shows the PRC is competing directly in decision systems.

2024 activity reported in 2025 · Moderate confidence

The assessment does not publish unit adoption, decision metrics, or combat results.

AI-O11 · U.S. intended use

The classified-network announcement states intended benefits in data synthesis, situational understanding, and warfighter decision support, but reports no measured baseline or achieved effect.

Decision-performance evidence

A stated benefit or plausible use case remains an objective until a controlled comparison shows time, quality, workload, or mission improvement.

Outcome not yet reported

Prevents model access from being scored as decision superiority.

May 2026 · High confidence

Operational results may be classified or collected after the announcement.

AI-O12 · Comparative mission outcome

No public dataset compares U.S./allied and PLA human-plus-AI decision time, accuracy, workload, error severity, and mission effect on representative national-security tasks.

Military decision advantage

When the defining outcome lacks comparable observations, it remains unknown rather than inherited from model, compute, or publication leadership.

Unknown comparison

Keeps operational decision performance contested.

August 2026 · High confidence

The absence is a public-evidence gap, not proof that effects do not exist.

05 · AI-T5
Assurance and contested robustnessReliable behavior after deployment and under adversary pressure
ContestedLow–moderate confidence

Mission effectMaintain useful, secure, calibrated AI behavior as models, data, users, networks, and threats change.

Analytical scopeEvaluation validity, cyber and misuse risk, adversarial inputs, drift, disconnected operation, human override, monitoring, and recovery.

Action priorityTurn evaluation into an operational control loop

  1. 01Pre-deployment test
  2. 02Security & red team
  3. 03Live monitoring
  4. 04Rollback & recovery
Judgment

The U.S. measurement ecosystem is a competitive asset, but NIST's own findings show benchmark and post-deployment monitoring practice is still immature. There is no demonstrated comparative field-robustness lead.

What offsets it

Independent held-out evaluation, provider diversity, logging, red teaming, rollback, edge testing, and continuous mission monitoring reduce dependence on one benchmark or model.

What remains unknown

Classified incident rates, adversarial success, drift, edge performance, operator misuse, and mission recovery are not publicly comparable.

Observation → applied rule → condition → driver effect3 observations
AI-O13 · Evaluation validity

NIST AI 800-2 describes current language-model and agent benchmark practices as preliminary and organizes them around defining the measurement target, implementing the evaluation, and analyzing/reporting results.

Reproducible model evaluation

A transparent evaluation process improves decision quality, but preliminary voluntary practice is an enabler rather than field assurance.

Evaluation capability maturing

Provides a modest U.S. institutional advantage.

January 2026 draft · High confidence

The document is an initial public draft and automated benchmarks cover only part of assurance.

AI-O14 · Independent versus self-reported results

CAISI found DeepSeek V4 appeared roughly at frontier parity on developer-reported benchmarks but lagged on CAISI's precommitted suite, including held-out reasoning, software, and cyber tasks.

Evaluation independence

A material difference between developer-selected and independent held-out results requires procurement decisions to use independent, mission-specific evaluation.

Selection bias demonstrated

Validates independent testing as a competitive control, not paperwork.

April 2026 · High confidence

No benchmark suite is neutral to all scaffolding and task choices.

AI-O15 · Post-deployment monitoring

NIST AI 800-4 concludes that monitoring is crucial for reliability, unexpected behavior, and real-world consequences, while common methods, terminology, guidance, and information sharing remain nascent and fragmented.

6 monitoring categories
Live AI assurance

If validated monitoring practice is immature, pre-deployment benchmark success cannot be treated as sustained operational reliability.

Assurance gap

Keeps contested robustness from becoming a U.S. edge.

March 2026 · High confidence

The report identifies challenges and categories rather than grading specific military deployments.

Improve the advantage

4 actions tied to diagnosed nodes.

Lead time is an implementation attribute—not a forecasted future rating. Every action has an owner, prerequisite, and observable completion test.

AI-A101
Recommendation

Create recurring independent mission-model trials

Frontier rankings and provider claims do not show which model, scaffold, data, and operator combination improves a specific military decision.

Maintain held-out mission suites spanning planning, cyber, intelligence, logistics, autonomy, and contested edge use; evaluate multiple U.S. and adversary models with controlled tools, token budgets, latency, cost, uncertainty, and human baselines.

Implementation test

PrerequisitesProtected datasets, independent evaluators, reproducible harnesses, provider access, and releasable summary methods.

Verify successEvery priority deployment has current independent results against a human and non-AI baseline, including cost, latency, failure, and adversarial conditions.

Linked threadsAI-T1 · AI-T3 · AI-T4 · AI-T5

Expected effectPreserves the model-quality edge, detects PRC convergence, and makes procurement follow measured mission value.

FeasibilityHigh

Cost bandMedium

Lead time6–18 months

OwnerDOD Chief Digital and AI Office, CAISI, services and combatant commands

AI-A202
Recommendation

Convert classified access into instrumented operator adoption

Eight-provider availability can remain shelfware if mission data, tools, permissions, training, and workflow ownership are missing.

Select consequential workflows; integrate data and tools; deploy in shadow mode; measure decision time, quality, workload, override, and downstream effect; scale only where a named owner accepts the result.

Implementation test

PrerequisitesIL6/IL7 access, mission-data contracts, workflow owners, operator training, and evaluation/monitoring pipelines.

Verify successAdopted systems deliver statistically and operationally meaningful improvement against baseline without unacceptable error or hidden manual burden.

Linked threadsAI-T3 · AI-T4

Expected effectTurns commercial frontier strength into operational decision advantage while preserving competition and human accountability.

FeasibilityHigh

Cost bandMedium–high

Lead time1–3 years

OwnerDOD CDAO, combatant commands and services

AI-A303
Recommendation

Protect the full-stack ecosystem and widen allied access

Model leadership depends on compute, energy, chips, cloud, talent, data, tools, and high-impact research—not one export-control chokepoint.

Expand power and data-center delivery, allied semiconductor and cloud capacity, research-compute access, talent attraction, secure full-stack export packages, and rapid substitution plans while testing whether controls produce the intended capability effect.

Implementation test

PrerequisitesInfrastructure permits, energy supply, security agreements, research funding, immigration/talent policy, and outcome-based control review.

Verify successAccessible allied compute, model, and research capacity grows while critical single-source exposure and time to add power or replace constrained components decline.

Linked threadsAI-T2

Expected effectSustains frontier iteration, diffuses U.S. standards and platforms, and reduces single-region or single-provider dependence.

FeasibilityModerate

Cost bandVery high

Lead time2–8 years

OwnerCommerce, Energy, State, DOD, NSF, industry and allies

AI-A404
Recommendation

Make live monitoring, rollback, and red teaming part of the weapon system

AI behavior changes with model updates, data, prompts, users, adversary inputs, and infrastructure; pre-deployment approval decays quickly.

Require mission-specific telemetry, drift and incident detection, adversarial testing, user feedback, human override, safe degradation, model/version provenance, rollback, and decommissioning for every operational AI service.

Implementation test

PrerequisitesLogging standards, evaluation triggers, protected incident sharing, model inventory, and operational rollback authority.

Verify successDeployments detect representative drift and attack, revert safely, and restore approved performance within mission thresholds during exercises and live incidents.

Linked threadsAI-T5

Expected effectPreserves trust and mission continuity while allowing faster iteration than one-time certification.

FeasibilityModerate–high

Cost bandMedium

Lead time1–3 years

OwnerDOD CDAO, test organizations, authorizing officials and system owners

Analytical audit trail

The machinery is available—but subordinate.

The finding comes first. Open this section to inspect composition rules, evidence limitations, and every source record.

Methods & evidenceInspect how the current judgment was reached
Scope rule

Current evidence only.

The assessment evaluates AI capability and decision advantage usable now using public evidence through August 5, 2026. Model quality, enabling inputs, access, adoption, mission performance, and assurance are separate dimensions.

Selection rule

Consequence before convenience.

Threads are included when they represent an independently varying source of advantage or failure: model capability/efficiency, ecosystem capacity, secure access, mission outcome, and robustness.

Composition rule

No black-box average.

Direct independent performance and observed operational outcome outrank compute, publications, contracts, or announced use cases. The overall score is adjudicated and not a weighted input index.

Evidence rule

Observation, inference, and unknown stay separate.

Held-out government evaluation, official infrastructure and adoption records, harmonized science indicators, and measurement-science findings are separated from developer claims and strategic intent.

Classified models, mission datasets, user adoption, decision outcomes, incidents, and PLA performance remain unknown. Missing outcomes are not inferred from inputs on either side.

Known limitations

What this public assessment cannot prove.

  • Frontier model rankings change quickly and depend on benchmark, scaffolding, token budget, serving configuration, and model availability.
  • Public benchmarks do not represent all military decision, multimodal, autonomy, deception, and disconnected-edge tasks.
  • Compute, publication, patent, and doctorate indicators are enabling-capacity proxies rather than operational AI effects.
  • Classified U.S. and PLA adoption, mission outcomes, model incidents, and human-machine performance prevent symmetric operational comparison.
  • Provider agreements and advertised AI features are credited as access or intent until repeatable mission output is observed.
Public evidence

9 source records; 9 primary or direct records.

AI-S1 · Tier 1CAISI Evaluation of DeepSeek V4 Pro

Primary use: Independent U.S./PRC frontier capability, task-level benchmark, cost, and evaluation-selection evidence.

Known limitation: Bounded public-model comparison whose results depend on benchmark suite, scaffolding, and serving choices.

Open public source ↗
AI-S2 · Tier 1CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks

Primary use: Earlier cross-model capability, security, censorship, adoption, and cost evidence.

Known limitation: Superseded in part by newer model releases; remains useful for evaluation continuity and security dimensions.

Open public source ↗
AI-S3 · Tier 1Classified Networks AI Agreements

Primary use: Eight-provider IL6/IL7 access, intended use, and classified-network deployment pathway.

Known limitation: Official announcement does not report adoption, uptime, cost, model versions, or achieved mission effects.

Open public source ↗
AI-S4 · Tier 12025 Military and Security Developments Involving the PRC

Primary use: PRC military AI applications, model progress, accelerator constraints, military-civil fusion, doctrine, and experimentation.

Known limitation: Unclassified threat assessment; PLA deployment, performance, and underlying methods are partly undisclosed.

Open public source ↗
AI-S5 · Tier 1The State of U.S. Science and Engineering 2026

Primary use: Comparable R&D, publication, citation, patent, collaboration, and science-workforce indicators.

Known limitation: Broad and often lagged science indicators are not direct measures of frontier model or military performance.

Open public source ↗
AI-S6 · Tier 1NAIRR at Two Years: Advancing American AI Innovation and Leadership

Primary use: Research-team, student, geographic, compute, data, and public-private access evidence.

Known limitation: Participation and resource access do not establish research quality, model leadership, or military relevance.

Open public source ↗
AI-S7 · Tier 1NIST AI 800-2: Practices for Automated Benchmark Evaluations

Primary use: Evaluation objectives, implementation, analysis, reporting, and preliminary-practice limits.

Known limitation: Initial public draft focused on automated benchmarks, not complete operational assurance.

Open public source ↗
AI-S8 · Tier 1NIST AI 800-4: Challenges to Monitoring Deployed AI Systems

Primary use: Post-deployment monitoring importance, categories, gaps, barriers, and immature practice.

Known limitation: Cross-sector research identifies challenges but does not evaluate named military deployments.

Open public source ↗
AI-S9 · Tier 1Artificial Intelligence: Key Practices to Help Ensure Accountability in Federal Use

Primary use: Governance, data, performance, and monitoring accountability framework for operational adoption.

Known limitation: Federal accountability framework predates the latest frontier models and is not a comparative U.S./PRC performance assessment.

Open public source ↗
Structured current assessment

Reuse the evidence—not just the conclusion.

The public JSON contains the executive judgment, drivers, decisive threads, observations, actions, limitations, and source records.

Assessment boundary

More traceability, not false precision.

This applied layer substantiates the current call without projecting future ratings or silently changing the core ten-area dataset. Compare the core capability record →