{"schema_version":"1.0","slug":"ai-decision","number":6,"title":"AI & decision","evidence_cutoff":"August 5, 2026","overall":{"score":0.75,"label":"U.S. edge","confidence":"Moderate","judgment":"The United States retains the current AI edge because independently tested frontier models are more capable, advanced compute and cloud ecosystems are deeper, and eight leading providers are being connected to classified military networks. The edge is narrower than compute or investment headlines imply: PRC models are improving quickly and can be cheaper at similar capability, China's research and talent scale is enormous, and neither side has publicly demonstrated a broad, repeatable military decision-performance advantage under realistic attack and data constraints.","composition_note":"Model quality is measured directly and kept separate from enabling inputs. NIST's independent evaluation places the leading tested PRC model about eight months behind the U.S. frontier, while also finding it cheaper than a comparable U.S. reference on five of seven cost tests. Compute, research, access, operational adoption, mission outcome, and assurance are therefore adjudicated independently. Classified-network agreements count as access—not decision superiority—and announced AI features count only when operational effect is observed."},"takeaways":[{"eyebrow":"Strongest U.S. advantage","title":"Frontier model capability","body":"Independent NIST testing finds a measurable U.S. lead across cyber, software engineering, science, reasoning, and mathematics rather than inferring quality from compute alone."},{"eyebrow":"Most important PRC offset","title":"Efficiency and rapid diffusion","body":"DeepSeek V4 was cheaper than a similar-capability U.S. reference on five of seven tested workloads, making deployment economics a separate competitive dimension."},{"eyebrow":"Largest analytical gap","title":"Military decision performance","body":"Public evidence does not show which side consistently improves decision time, accuracy, workload, or mission effect with AI in representative contested operations."},{"eyebrow":"Decision","title":"Buy measured mission improvement","body":"Pair access to competing models with mission datasets, operator trials, security testing, and post-deployment monitoring; pay for repeatable outcome gains, not tokens or demonstrations."}],"drivers":[{"id":"AI-D1","name":"Model quality & demonstrated capability","score":1.5,"label":"U.S. lead","confidence":"High","judgment":"Independent CAISI testing places the strongest evaluated PRC model about eight months behind the U.S. frontier across five capability domains.","contrary":"The lag is an estimate from a bounded benchmark suite, not a universal intelligence measure; PRC models perform near the U.S. frontier on several individual tasks.","source_ids":["AI-S1","AI-S2"]},{"id":"AI-D2","name":"Compute, chips & cloud infrastructure","score":1.25,"label":"U.S. edge","confidence":"Moderate","judgment":"The U.S. and accessible allies control the deeper frontier accelerator, cloud, software, and model-provider stack, while the PRC remains constrained in access to leading accelerators.","contrary":"PRC firms improve hardware/software efficiency, stockpile components, pursue domestic alternatives, and use intermediaries; aggregate U.S. compute does not equal military access at the edge.","source_ids":["AI-S3","AI-S4"]},{"id":"AI-D3","name":"Cost, efficiency & diffusion","score":-0.5,"label":"PRC edge","confidence":"Moderate","judgment":"DeepSeek V4 demonstrates that a PRC model can deliver similar aggregate capability at lower end-to-end cost on most tested workloads, supporting faster diffusion despite a frontier-quality lag.","contrary":"The cost relationship varied by benchmark and excluded two evaluations; developer pricing, hosting configuration, reliability, and security can change total mission cost.","source_ids":["AI-S1"]},{"id":"AI-D4","name":"Talent, research & data ecosystem","score":0.5,"label":"U.S. edge","confidence":"Low–moderate","judgment":"The United States retains disproportionate high-impact science, patents, global collaboration, and private-sector translation, while China leads publication volume, doctorate output, and several patent counts.","contrary":"Broad science indicators are imperfect AI proxies; China surpassed the United States in comparable R&D volume and is building a formidable talent and innovation base.","source_ids":["AI-S5","AI-S6"]},{"id":"AI-D5","name":"Secure model access & deployment pathways","score":1.0,"label":"U.S. edge","confidence":"Moderate","judgment":"Agreements with eight frontier AI companies for IL6 and IL7 networks create a uniquely broad, competitive pathway to classified operational use.","contrary":"Agreements and network availability are not user adoption, mission integration, uptime, or verified effect; PLA military-civil fusion may move commercial advances through different channels.","source_ids":["AI-S3","AI-S4"]},{"id":"AI-D6","name":"Operational integration & decision performance","score":0.0,"label":"Contested","confidence":"Low","judgment":"Both militaries are integrating AI into analysis, cyber, unmanned systems, and planning, but public evidence does not establish a comparative, mission-specific improvement in decisions or outcomes.","contrary":"Classified adoption may be substantial; PLA doctrine and U.S. provider access show serious institutional commitment even where outcomes remain undisclosed.","source_ids":["AI-S3","AI-S4"]},{"id":"AI-D7","name":"Evaluation, assurance & contested robustness","score":0.25,"label":"Contested","confidence":"Low–moderate","judgment":"The United States has a stronger public evaluation and measurement-science ecosystem, but benchmark validity, deployment monitoring, cyber resilience, edge performance, and human-machine failure remain immature.","contrary":"NIST's own work emphasizes that common evaluation and monitoring practices are preliminary and fragmented; a strong standards ecosystem does not prove fielded systems are safe or robust.","source_ids":["AI-S1","AI-S7","AI-S8","AI-S9"]}],"threads_heading":"Five parallel AI threads determine usable decision advantage.","threads_intro":"There is no single compute pipeline. Model capability, efficiency, research and infrastructure, secure military access, mission outcome, and assurance can move independently. Each thread shows which evidence is a direct outcome and which is only an enabling input.","threads":[{"id":"AI-T1","name":"Frontier model capability and efficiency","decisive_node":"Independent task performance at mission-relevant cost","mission_effect":"Provide higher-quality reasoning, coding, cyber, scientific, perception, and agentic assistance than the opposing ecosystem.","scope":"Publicly released models evaluated under controlled scaffolding and token budgets; classified models and undocumented deployments remain unknown.","priority":"Preserve and measure the edge","score":1.25,"label":"U.S. edge","confidence":"High","judgment":"U.S. frontier quality leads in the strongest available independent comparison, but PRC efficiency materially narrows the operational and diffusion advantage.","mitigation":"Competition across providers, open-weight models, inference optimization, and workload-specific routing can lower U.S. cost without abandoning higher-capability options.","gap":"No benchmark suite fully represents military judgment, multimodal sensing, classified data, deception, human teaming, or long-horizon operations.","nodes":[{"name":"Capability","status":"us"},{"name":"Cost efficiency","status":"prc"},{"name":"Mission fit","status":"unknown"},{"name":"Operational value","status":"unknown"}],"source_ids":["AI-S1","AI-S2","AI-S7"],"recommendation_ids":["AI-A1"],"observations":[{"id":"AI-O01","driver_ids":["AI-D1"],"node":"Aggregate capability","observation":"CAISI's April 2026 evaluation estimated that DeepSeek V4, the strongest PRC model it had tested, lagged the leading U.S. frontier by about eight months across cyber, software engineering, natural sciences, abstract reasoning, and mathematics.","value":"~8","unit":"months estimated frontier lag","metric":"Independent frontier capability","applied_rule":"A precommitted, multi-domain independent evaluation with held-out tasks receives greater weight than self-reported benchmarks or input proxies.","bucket":"U.S. frontier lead","effect":"Directly supports a U.S. lead in model quality.","as_of":"April 2026 evaluation","confidence":"High","limitation":"The time-lag estimate depends on selected models, benchmarks, scaffolding, and a fitted trend.","source_ids":["AI-S1"]},{"id":"AI-O02","driver_ids":["AI-D1"],"node":"Task-level performance","observation":"In CAISI's table, the leading U.S. model scored 71 percent versus DeepSeek V4's estimated 32 percent on a cyber benchmark and 78 versus 44 percent on the held-out PortBench software task, while results were close on several science and mathematics tests.","value":"71–32 / 78–44","unit":"percentage-point benchmark comparisons","metric":"Domain-specific demonstrated capability","applied_rule":"Large differences on held-out, mission-relevant task families establish a direct capability edge; near-parity elsewhere prevents universalizing it.","bucket":"U.S. edge with domain variation","effect":"Shows where the model lead is consequential and where it is not.","as_of":"April 2026 evaluation","confidence":"High","limitation":"Benchmarks are proxies and the cyber value for DeepSeek was imputed from a subset of samples.","source_ids":["AI-S1"]},{"id":"AI-O03","driver_ids":["AI-D3"],"node":"Cost efficiency","observation":"DeepSeek V4 was less expensive than the selected similar-capability U.S. reference on five of seven comparable benchmarks; its relative cost ranged from 53 percent lower to 41 percent higher.","value":"5 of 7","unit":"benchmarks with lower DeepSeek cost","metric":"End-to-end cost for correctly solved tasks","applied_rule":"Workload-level cost per successful outcome is a distinct competitive measure and cannot be inferred from token price alone.","bucket":"PRC efficiency offset","effect":"Reduces the practical size of the U.S. capability lead.","as_of":"April 2026 evaluation","confidence":"High","limitation":"Two benchmarks were excluded and provider prices, hardware, and serving configurations can change.","source_ids":["AI-S1"]}]},{"id":"AI-T2","name":"Compute, talent and research ecosystem","decisive_node":"Sustain frontier development and broad adoption","mission_effect":"Train, serve, improve, and diffuse capable AI systems while replacing constrained hardware, energy, data, and talent inputs.","scope":"Advanced accelerators, cloud, R&D, high-impact output, talent, public research access, and substitution—not raw compute or publication count alone.","priority":"Keep the ecosystem open and resilient","score":0.75,"label":"U.S. edge","confidence":"Moderate","judgment":"The U.S. ecosystem retains frontier hardware, highly cited research, and business-led translation advantages, while China's scale in R&D, publications, doctorates, patents, and efficiency makes the contest much closer than a chip-control narrative suggests.","mitigation":"Allied semiconductor depth, NAIRR access, energy and data-center expansion, talent attraction, and research collaboration expand the accessible U.S. base.","gap":"Comparable frontier training compute, usable military data, researcher flows, energy constraints, and domestic chip yield are incomplete or proprietary.","nodes":[{"name":"Compute & energy","status":"us"},{"name":"Talent & research","status":"mixed"},{"name":"Data & tools","status":"mixed"},{"name":"Frontier iteration","status":"us"}],"source_ids":["AI-S4","AI-S5","AI-S6"],"recommendation_ids":["AI-A3"],"observations":[{"id":"AI-O04","driver_ids":["AI-D2"],"node":"Advanced accelerators","observation":"The 2025 Defense Department report assesses that China's AI sector remains constrained by limited access to high-performance accelerators while pursuing efficiency, stockpiles, circumvention, and domestic alternatives.","value":"","unit":"","metric":"Frontier compute access","applied_rule":"A documented chokepoint supports a relative edge only when the competitor's active substitutions and observed model performance are also considered.","bucket":"PRC compute constraint","effect":"Supports a U.S. compute edge but not a complete AI lead.","as_of":"2024 activity reported in 2025","confidence":"Moderate–high","limitation":"Stockpiles, illicit access, domestic output, and actual training allocations are uncertain.","source_ids":["AI-S4"]},{"id":"AI-O05","driver_ids":["AI-D4"],"node":"Research scale and impact","observation":"NSF's 2026 indicators estimate China produced 31 percent of global science and engineering articles in 2024 versus 12 percent for the United States, while U.S. work retained a disproportionate share of highly cited output and internationally coauthored articles.","value":"31 / 12","unit":"percent of global S&E articles","metric":"Research scale and impact","applied_rule":"Volume, citation impact, and collaboration are separate measures; high volume cannot be equated with frontier quality, and high impact cannot erase competitor scale.","bucket":"Split research advantage","effect":"Holds the talent and research driver to a narrow U.S. edge.","as_of":"2024 output reported in 2026","confidence":"High","limitation":"All-science publications are a broad AI ecosystem proxy and citation measures lag current work.","source_ids":["AI-S5"]},{"id":"AI-O06","driver_ids":["AI-D4"],"node":"Research access","observation":"The NSF-led NAIRR reported supporting more than 600 research teams and 6,000 students across all U.S. states, Washington, D.C., and Puerto Rico through public-private compute, data, models, and expertise.","value":"600 / 6000","unit":"research teams / students","metric":"Distributed research-resource access","applied_rule":"Observed access to advanced resources and training is a current ecosystem enabler, not direct model or military performance.","bucket":"U.S. diffusion strength","effect":"Broadens the U.S. innovation base beyond frontier firms.","as_of":"March 2026","confidence":"High","limitation":"Participation does not establish research quality, military relevance, or comparison with PRC programs.","source_ids":["AI-S6"]}]},{"id":"AI-T3","name":"Secure government access and adoption","decisive_node":"Operators can use competing models on mission networks","mission_effect":"Move frontier and specialized AI into classified workflows with data, permissions, tools, and accountable human use.","scope":"Available providers, network levels, data connections, user access, latency, training, acquisition, and sustained use—not contract count.","priority":"Convert access into adoption","score":0.75,"label":"U.S. edge","confidence":"Moderate","judgment":"The United States has created a broad classified-access pathway across eight leading firms, but public evidence does not yet show user scale, mission availability, or operational effect.","mitigation":"Provider competition, portable evaluations, shared security services, and mission-level procurement can avoid lock-in and accelerate adoption.","gap":"Active users, approved use cases, model versions, uptime, data access, cost, and field effects on IL6/IL7 networks are not public.","nodes":[{"name":"Provider access","status":"us"},{"name":"Mission data","status":"unknown"},{"name":"User workflow","status":"mixed"},{"name":"Sustained adoption","status":"unknown"}],"source_ids":["AI-S3","AI-S4"],"recommendation_ids":["AI-A1","AI-A2"],"observations":[{"id":"AI-O07","driver_ids":["AI-D5"],"node":"Classified access","observation":"In May 2026 the Department announced agreements with eight leading AI companies to make advanced capabilities available on Impact Level 6 and 7 classified networks.","value":"8","unit":"frontier AI companies","metric":"Competitive classified model access","applied_rule":"Multiple providers on mission networks create a current access and competition advantage, while adoption and effect remain separate thresholds.","bucket":"Broad U.S. access pathway","effect":"Supports a U.S. edge in secure deployment options.","as_of":"May 2026","confidence":"High","limitation":"The release does not report operational availability, user counts, model versions, cost, or mission outcome.","source_ids":["AI-S3"]},{"id":"AI-O08","driver_ids":["AI-D5","AI-D6"],"node":"PLA commercial integration","observation":"The 2025 PRC report assesses that military-civil fusion gives the PLA continuing access to commercial and academic AI advances and records AI-related work in unmanned systems, ISR analysis, decision assistance, cyber operations, and information campaigns.","value":"","unit":"","metric":"Commercial-to-military diffusion","applied_rule":"A structured diffusion mechanism and observed applications establish serious adoption activity, but not comparative mission effect.","bucket":"PRC adoption pathway","effect":"Prevents the U.S. access advantage from becoming a lead.","as_of":"2024 activity reported in 2025","confidence":"Moderate","limitation":"Specific fielded systems, users, performance, and security are not disclosed.","source_ids":["AI-S4"]},{"id":"AI-O09","driver_ids":["AI-D5","AI-D6"],"node":"Sustained operational adoption","observation":"No public source establishes the share of U.S. or PLA operational units that routinely use AI for priority decisions, the tasks delegated, or the resulting performance.","value":"","unit":"","metric":"Adoption at operational scale","applied_rule":"Network availability, contracts, demonstrations, and doctrine cannot substitute for sustained use by trained operators in real workflows.","bucket":"Unknown comparison","effect":"Caps confidence in the U.S. secure-access edge.","as_of":"August 2026","confidence":"High","limitation":"Classified adoption may be substantial on both sides.","source_ids":["AI-S3","AI-S4"]}]},{"id":"AI-T4","name":"Mission decision performance","decisive_node":"Observed improvement in a defined operational decision","mission_effect":"Reduce decision time or workload and improve accuracy, allocation, targeting, logistics, cyber response, or mission outcome without unacceptable error.","scope":"Context-specific human-plus-AI performance against a baseline, including latency, confidence, override, failure, and downstream effect.","priority":"Make the central claim measurable","score":0.0,"label":"Contested","confidence":"Low","judgment":"Public evidence supports strong capabilities and active adoption pathways, but neither side demonstrates a broad current advantage in military decision outcomes.","mitigation":"Instrumented operator trials, shadow mode, controlled rollout, and mission-level A/B comparisons can produce direct evidence without exposing classified tactics.","gap":"Relevant baselines, tasks, error costs, decision authority, adversary adaptation, and operational outcomes are classified or not standardized.","nodes":[{"name":"Mission task","status":"mixed"},{"name":"Human–AI workflow","status":"unknown"},{"name":"Decision quality","status":"unknown"},{"name":"Mission outcome","status":"unknown"}],"source_ids":["AI-S3","AI-S4","AI-S7"],"recommendation_ids":["AI-A1","AI-A2"],"observations":[{"id":"AI-O10","driver_ids":["AI-D6"],"node":"PLA decision concept","observation":"The PRC report describes Multi-Domain Precision Warfare as using big data and AI to aggregate information, identify weak points, and support rapid operational decisions, while stating that PLA intelligentized-warfare theory and concepts remain under development and experimentation.","value":"","unit":"","metric":"Operational concept maturity","applied_rule":"A coherent doctrine and recurring experimentation establish intent and learning, but not demonstrated decision advantage.","bucket":"Serious development / effect unproven","effect":"Shows the PRC is competing directly in decision systems.","as_of":"2024 activity reported in 2025","confidence":"Moderate","limitation":"The assessment does not publish unit adoption, decision metrics, or combat results.","source_ids":["AI-S4"]},{"id":"AI-O11","driver_ids":["AI-D6"],"node":"U.S. intended use","observation":"The classified-network announcement states intended benefits in data synthesis, situational understanding, and warfighter decision support, but reports no measured baseline or achieved effect.","value":"","unit":"","metric":"Decision-performance evidence","applied_rule":"A stated benefit or plausible use case remains an objective until a controlled comparison shows time, quality, workload, or mission improvement.","bucket":"Outcome not yet reported","effect":"Prevents model access from being scored as decision superiority.","as_of":"May 2026","confidence":"High","limitation":"Operational results may be classified or collected after the announcement.","source_ids":["AI-S3"]},{"id":"AI-O12","driver_ids":["AI-D6","AI-D7"],"node":"Comparative mission outcome","observation":"No public dataset compares U.S./allied and PLA human-plus-AI decision time, accuracy, workload, error severity, and mission effect on representative national-security tasks.","value":"","unit":"","metric":"Military decision advantage","applied_rule":"When the defining outcome lacks comparable observations, it remains unknown rather than inherited from model, compute, or publication leadership.","bucket":"Unknown comparison","effect":"Keeps operational decision performance contested.","as_of":"August 2026","confidence":"High","limitation":"The absence is a public-evidence gap, not proof that effects do not exist.","source_ids":["AI-S3","AI-S4","AI-S7"]}]},{"id":"AI-T5","name":"Assurance and contested robustness","decisive_node":"Reliable behavior after deployment and under adversary pressure","mission_effect":"Maintain useful, secure, calibrated AI behavior as models, data, users, networks, and threats change.","scope":"Evaluation validity, cyber and misuse risk, adversarial inputs, drift, disconnected operation, human override, monitoring, and recovery.","priority":"Turn evaluation into an operational control loop","score":0.25,"label":"Contested","confidence":"Low–moderate","judgment":"The U.S. measurement ecosystem is a competitive asset, but NIST's own findings show benchmark and post-deployment monitoring practice is still immature. There is no demonstrated comparative field-robustness lead.","mitigation":"Independent held-out evaluation, provider diversity, logging, red teaming, rollback, edge testing, and continuous mission monitoring reduce dependence on one benchmark or model.","gap":"Classified incident rates, adversarial success, drift, edge performance, operator misuse, and mission recovery are not publicly comparable.","nodes":[{"name":"Pre-deployment test","status":"us"},{"name":"Security & red team","status":"mixed"},{"name":"Live monitoring","status":"unknown"},{"name":"Rollback & recovery","status":"unknown"}],"source_ids":["AI-S1","AI-S2","AI-S7","AI-S8","AI-S9"],"recommendation_ids":["AI-A1","AI-A4"],"observations":[{"id":"AI-O13","driver_ids":["AI-D7"],"node":"Evaluation validity","observation":"NIST AI 800-2 describes current language-model and agent benchmark practices as preliminary and organizes them around defining the measurement target, implementing the evaluation, and analyzing/reporting results.","value":"","unit":"","metric":"Reproducible model evaluation","applied_rule":"A transparent evaluation process improves decision quality, but preliminary voluntary practice is an enabler rather than field assurance.","bucket":"Evaluation capability maturing","effect":"Provides a modest U.S. institutional advantage.","as_of":"January 2026 draft","confidence":"High","limitation":"The document is an initial public draft and automated benchmarks cover only part of assurance.","source_ids":["AI-S7"]},{"id":"AI-O14","driver_ids":["AI-D7"],"node":"Independent versus self-reported results","observation":"CAISI found DeepSeek V4 appeared roughly at frontier parity on developer-reported benchmarks but lagged on CAISI's precommitted suite, including held-out reasoning, software, and cyber tasks.","value":"","unit":"","metric":"Evaluation independence","applied_rule":"A material difference between developer-selected and independent held-out results requires procurement decisions to use independent, mission-specific evaluation.","bucket":"Selection bias demonstrated","effect":"Validates independent testing as a competitive control, not paperwork.","as_of":"April 2026","confidence":"High","limitation":"No benchmark suite is neutral to all scaffolding and task choices.","source_ids":["AI-S1"]},{"id":"AI-O15","driver_ids":["AI-D7"],"node":"Post-deployment monitoring","observation":"NIST AI 800-4 concludes that monitoring is crucial for reliability, unexpected behavior, and real-world consequences, while common methods, terminology, guidance, and information sharing remain nascent and fragmented.","value":"6","unit":"monitoring categories","metric":"Live AI assurance","applied_rule":"If validated monitoring practice is immature, pre-deployment benchmark success cannot be treated as sustained operational reliability.","bucket":"Assurance gap","effect":"Keeps contested robustness from becoming a U.S. edge.","as_of":"March 2026","confidence":"High","limitation":"The report identifies challenges and categories rather than grading specific military deployments.","source_ids":["AI-S8"]}]}],"recommendations":[{"id":"AI-A1","rank":1,"title":"Create recurring independent mission-model trials","diagnosis":"Frontier rankings and provider claims do not show which model, scaffold, data, and operator combination improves a specific military decision.","action":"Maintain held-out mission suites spanning planning, cyber, intelligence, logistics, autonomy, and contested edge use; evaluate multiple U.S. and adversary models with controlled tools, token budgets, latency, cost, uncertainty, and human baselines.","expected_effect":"Preserves the model-quality edge, detects PRC convergence, and makes procurement follow measured mission value.","feasibility":"High","cost_band":"Medium","lead_time":"6–18 months","owner":"DOD Chief Digital and AI Office, CAISI, services and combatant commands","prerequisites":"Protected datasets, independent evaluators, reproducible harnesses, provider access, and releasable summary methods.","verification":"Every priority deployment has current independent results against a human and non-AI baseline, including cost, latency, failure, and adversarial conditions.","thread_ids":["AI-T1","AI-T3","AI-T4","AI-T5"],"source_ids":["AI-S1","AI-S7"]},{"id":"AI-A2","rank":2,"title":"Convert classified access into instrumented operator adoption","diagnosis":"Eight-provider availability can remain shelfware if mission data, tools, permissions, training, and workflow ownership are missing.","action":"Select consequential workflows; integrate data and tools; deploy in shadow mode; measure decision time, quality, workload, override, and downstream effect; scale only where a named owner accepts the result.","expected_effect":"Turns commercial frontier strength into operational decision advantage while preserving competition and human accountability.","feasibility":"High","cost_band":"Medium–high","lead_time":"1–3 years","owner":"DOD CDAO, combatant commands and services","prerequisites":"IL6/IL7 access, mission-data contracts, workflow owners, operator training, and evaluation/monitoring pipelines.","verification":"Adopted systems deliver statistically and operationally meaningful improvement against baseline without unacceptable error or hidden manual burden.","thread_ids":["AI-T3","AI-T4"],"source_ids":["AI-S3","AI-S9"]},{"id":"AI-A3","rank":3,"title":"Protect the full-stack ecosystem and widen allied access","diagnosis":"Model leadership depends on compute, energy, chips, cloud, talent, data, tools, and high-impact research—not one export-control chokepoint.","action":"Expand power and data-center delivery, allied semiconductor and cloud capacity, research-compute access, talent attraction, secure full-stack export packages, and rapid substitution plans while testing whether controls produce the intended capability effect.","expected_effect":"Sustains frontier iteration, diffuses U.S. standards and platforms, and reduces single-region or single-provider dependence.","feasibility":"Moderate","cost_band":"Very high","lead_time":"2–8 years","owner":"Commerce, Energy, State, DOD, NSF, industry and allies","prerequisites":"Infrastructure permits, energy supply, security agreements, research funding, immigration/talent policy, and outcome-based control review.","verification":"Accessible allied compute, model, and research capacity grows while critical single-source exposure and time to add power or replace constrained components decline.","thread_ids":["AI-T2"],"source_ids":["AI-S4","AI-S5","AI-S6"]},{"id":"AI-A4","rank":4,"title":"Make live monitoring, rollback, and red teaming part of the weapon system","diagnosis":"AI behavior changes with model updates, data, prompts, users, adversary inputs, and infrastructure; pre-deployment approval decays quickly.","action":"Require mission-specific telemetry, drift and incident detection, adversarial testing, user feedback, human override, safe degradation, model/version provenance, rollback, and decommissioning for every operational AI service.","expected_effect":"Preserves trust and mission continuity while allowing faster iteration than one-time certification.","feasibility":"Moderate–high","cost_band":"Medium","lead_time":"1–3 years","owner":"DOD CDAO, test organizations, authorizing officials and system owners","prerequisites":"Logging standards, evaluation triggers, protected incident sharing, model inventory, and operational rollback authority.","verification":"Deployments detect representative drift and attack, revert safely, and restore approved performance within mission thresholds during exercises and live incidents.","thread_ids":["AI-T5"],"source_ids":["AI-S8","AI-S9"]}],"methodology":{"scope":"The assessment evaluates AI capability and decision advantage usable now using public evidence through August 5, 2026. Model quality, enabling inputs, access, adoption, mission performance, and assurance are separate dimensions.","selection_rule":"Threads are included when they represent an independently varying source of advantage or failure: model capability/efficiency, ecosystem capacity, secure access, mission outcome, and robustness.","composition_rule":"Direct independent performance and observed operational outcome outrank compute, publications, contracts, or announced use cases. The overall score is adjudicated and not a weighted input index.","evidence_rule":"Held-out government evaluation, official infrastructure and adoption records, harmonized science indicators, and measurement-science findings are separated from developer claims and strategic intent.","unknown_rule":"Classified models, mission datasets, user adoption, decision outcomes, incidents, and PLA performance remain unknown. Missing outcomes are not inferred from inputs on either side."},"limitations":["Frontier model rankings change quickly and depend on benchmark, scaffolding, token budget, serving configuration, and model availability.","Public benchmarks do not represent all military decision, multimodal, autonomy, deception, and disconnected-edge tasks.","Compute, publication, patent, and doctorate indicators are enabling-capacity proxies rather than operational AI effects.","Classified U.S. and PLA adoption, mission outcomes, model incidents, and human-machine performance prevent symmetric operational comparison.","Provider agreements and advertised AI features are credited as access or intent until repeatable mission output is observed."],"sources":[{"id":"AI-S1","title":"CAISI Evaluation of DeepSeek V4 Pro","url":"https://www.nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro","publisher":"National Institute of Standards and Technology","date":"May 2026","evidence_tier":"Tier 1","primary_use":"Independent U.S./PRC frontier capability, task-level benchmark, cost, and evaluation-selection evidence.","limitation":"Bounded public-model comparison whose results depend on benchmark suite, scaffolding, and serving choices."},{"id":"AI-S2","title":"CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks","url":"https://www.nist.gov/news-events/news/2025/09/caisi-evaluation-deepseek-ai-models-finds-shortcomings-and-risks","publisher":"National Institute of Standards and Technology","date":"September 2025","evidence_tier":"Tier 1","primary_use":"Earlier cross-model capability, security, censorship, adoption, and cost evidence.","limitation":"Superseded in part by newer model releases; remains useful for evaluation continuity and security dimensions."},{"id":"AI-S3","title":"Classified Networks AI Agreements","url":"https://www.war.gov/News/Releases/Release/Article/4475177/classified-networks-ai-agreements/","publisher":"U.S. Department of War","date":"May 2026","evidence_tier":"Tier 1","primary_use":"Eight-provider IL6/IL7 access, intended use, and classified-network deployment pathway.","limitation":"Official announcement does not report adoption, uptime, cost, model versions, or achieved mission effects."},{"id":"AI-S4","title":"2025 Military and Security Developments Involving the PRC","url":"https://media.defense.gov/2025/Dec/23/2003849070/-1/-1/1/ANNUAL-REPORT-TO-CONGRESS-MILITARY-AND-SECURITY-DEVELOPMENTS-INVOLVING-THE-PEOPLES-REPUBLIC-OF-CHINA-2025.PDF","publisher":"U.S. Department of Defense","date":"December 2025","evidence_tier":"Tier 1","primary_use":"PRC military AI applications, model progress, accelerator constraints, military-civil fusion, doctrine, and experimentation.","limitation":"Unclassified threat assessment; PLA deployment, performance, and underlying methods are partly undisclosed."},{"id":"AI-S5","title":"The State of U.S. Science and Engineering 2026","url":"https://ncses.nsf.gov/pubs/nsbsep20261/executive-summary","publisher":"National Science Board / National Science Foundation","date":"May 2026","evidence_tier":"Tier 1","primary_use":"Comparable R&D, publication, citation, patent, collaboration, and science-workforce indicators.","limitation":"Broad and often lagged science indicators are not direct measures of frontier model or military performance."},{"id":"AI-S6","title":"NAIRR at Two Years: Advancing American AI Innovation and Leadership","url":"https://www.nsf.gov/cise/updates/nairr-2-years-advancing-american-artificial-intelligence","publisher":"U.S. National Science Foundation","date":"March 2026","evidence_tier":"Tier 1","primary_use":"Research-team, student, geographic, compute, data, and public-private access evidence.","limitation":"Participation and resource access do not establish research quality, model leadership, or military relevance."},{"id":"AI-S7","title":"NIST AI 800-2: Practices for Automated Benchmark Evaluations","url":"https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf","publisher":"National Institute of Standards and Technology","date":"January 2026","evidence_tier":"Tier 1","primary_use":"Evaluation objectives, implementation, analysis, reporting, and preliminary-practice limits.","limitation":"Initial public draft focused on automated benchmarks, not complete operational assurance."},{"id":"AI-S8","title":"NIST AI 800-4: Challenges to Monitoring Deployed AI Systems","url":"https://www.nist.gov/publications/challenges-monitoring-deployed-ai-systems-center-ai-standards-and-innovation","publisher":"National Institute of Standards and Technology","date":"March 2026","evidence_tier":"Tier 1","primary_use":"Post-deployment monitoring importance, categories, gaps, barriers, and immature practice.","limitation":"Cross-sector research identifies challenges but does not evaluate named military deployments."},{"id":"AI-S9","title":"Artificial Intelligence: Key Practices to Help Ensure Accountability in Federal Use","url":"https://www.gao.gov/products/gao-23-106811","publisher":"U.S. Government Accountability Office","date":"September 2023","evidence_tier":"Tier 1","primary_use":"Governance, data, performance, and monitoring accountability framework for operational adoption.","limitation":"Federal accountability framework predates the latest frontier models and is not a comparative U.S./PRC performance assessment."}]}