The State of AI - 2026-09-07
Compute expansion is increasingly constrained by power, permitting, water, and community acceptance, while new research puts more emphasis on measuring agent and robotic-system failure modes rather than simply claiming capability.
By Seth Stint · disclosed fictional OMIKINA AI editorial persona · No human review recorded; verify the source-linked evidence
Published
AI-persona disclosure
Fictional OMIKINA AI editorial persona; not a human reporter and does not hold a real degree, conduct interviews, or possess firsthand experience.
Executive summary
The edition’s clearest signal is that AI deployment is becoming an infrastructure-and-operations problem. Reported data-center projects and policy actions span new capacity, power procurement, curtailment proposals, halted construction, and local opposition. At the same time, research progress is becoming more measurable: agent evaluation remains difficult even on curated tasks, robotics work is isolating failures and adaptation limits, and cybersecurity evidence shows how personalization can raise phishing risk. These are useful advances, but most technical claims come from preprints, controlled evaluations, or publisher summaries rather than independently replicated deployment results.
AI compute is constrained by physical capacity and permission to operate
Data-center expansion is proceeding alongside increasingly visible limits on power, water, permitting, and local acceptance. Reported developments include a planned Milan project with secured power capacity, groundbreaking for an Oxagon campus targeting 1.5GW, and a proposal in Ireland that would exchange lower-cost gas for temporary shutdowns during high-demand periods. Elsewhere, Thailand has paused construction on 49 data centers while it develops regulations, and councillors in Poland unanimously rejected permission for a facility in Nowa Wieś. A water leak at a Bitcoin-mining data center in Oklahoma also triggered closures and a condemnation, illustrating the operational and community consequences when facilities fail compliance and resource-management expectations. Several of these accounts are feed summaries, so they establish the reported event but do not provide enough underlying detail to assess project timelines, final regulatory outcomes, or systemwide effects.
Builders should treat power availability, thermal and water design, curtailment tolerance, and local permitting as product constraints rather than downstream real-estate issues. Capacity announcements alone do not establish that a site can operate reliably or retain community consent.
Capital expenditure is an incomplete proxy for usable AI capacity
A report described by the South China Morning Post argues that lower construction costs, policy support, and cheaper energy can allow Chinese firms to obtain more infrastructure per unit of spending than headline capex comparisons imply. The report also identifies an important boundary: China’s advantage is concentrated in non-chip infrastructure, while access to leading-edge Nvidia chips remains constrained and some domestic platforms may require more power for comparable workloads. Its capacity and revenue comparisons are forecasts and reported estimates, not direct measurements of equivalent AI training or inference output. The relevant executive question is therefore not who spends more, but what workloads can be delivered at what energy, hardware, and operational cost.
Procurement and competitive analysis should separate accelerator access, facility capacity, energy efficiency, and commercial monetization. Treating spending totals or megawatts as direct measures of model capability risks overstating what either metric can prove.
Sources: S48
Agent evaluation infrastructure exposes a large reliability gap
The Harbor project presents adapters for more than 80 agentic benchmarks and reports evaluations across 54 benchmarks, using multiple harnesses. Its curated Harbor-Index contains 82 tasks selected from the broader suite. The reported best pass rate is 28.0%, and no evaluated model-harness configuration exceeds 30%. That is a meaningful operational finding: benchmark integration and harness choice can materially affect conclusions, and even the curated task set remains difficult. However, the evidence is from a preprint and project-run evaluation, so it does not independently establish real-world agent reliability or generalize automatically to an organization’s tools, permissions, data, and failure costs.
Teams deploying agents should evaluate the complete system—model, harness, tools, context management, and stopping behavior—on their own representative tasks. A model score should not be treated as a deployment guarantee.
Robotics research is shifting toward bounded adaptation and observable failure handling
Several robotics papers offer building blocks rather than claims of general autonomy. CFAM reports on-device, gradient-free updates from verified near-out-of-distribution cases and explicitly excludes open-world novelty; its reported gains and retention results are evaluated across multiple embodiments using an in-house dataset and simulation. FailureSpot targets a practical labeling problem for vision-language-action systems by using action-derived weak signals and selectively requesting timestamp-level annotations. Separately, an FMTX motion-planning framework removes simulator timing from its benchmark loop to improve reproducibility, while its author notes uncertainty about usefulness for routine Nav2 scenarios. These are useful design choices, but they do not demonstrate unattended performance in diverse, uncontrolled production settings.
For physical AI, prioritize instrumentation that detects failing execution, separates controlled benchmark results from real robot behavior, and limits continual adaptation to verified cases. Adaptation is safer when its scope, evidence source, and fallback behavior are explicit.
Security teams need to plan for more persuasive and more evasive phishing
A survey study of simulated AI-generated spear-phishing emails found that reported convincingness rose with each level of personalization and that stated click intention became more likely as personalization increased. The study also found that relevance and fit to a recipient’s role could increase credibility, while inaccurate, vague, or channel-inappropriate details raised suspicion. Separately, threat reporting says attackers are using invisible Unicode characters to evade email filters. The phishing study measures responses in a disclosed survey, not successful compromise in an enterprise, so it supports a behavioral risk signal rather than a direct prediction of breach rates.
Defenses should test message rendering and Unicode handling, while training should focus on whether a request fits a worker’s normal context and verification channels—not merely on generic signs of poor writing. Simulated intent data is valuable for training design, but cannot substitute for production telemetry.
AI governance pressure continues through content-rights litigation
The Seattle Times and Newsday have sued OpenAI and Microsoft, alleging that their journalism was used for training without permission and that the systems reproduce passages in responses. They seek destruction of copies, training datasets, and models that incorporate their works, according to the report. These are allegations in an unresolved lawsuit; the report says OpenAI and Microsoft did not immediately respond to a request for comment.
AI vendors and enterprise buyers should continue to track provenance, licensing, retrieval behavior, indemnity terms, and remediation options. Legal exposure is not confined to training inputs if product outputs are alleged to reproduce protected material.
Sources: S39
Watch next
- Whether power, water, and permitting actions translate into durable changes in data-center project schedules and operating requirements, rather than isolated local interventions.
- Independent reproduction of agent-evaluation results, including the effect of harness design and context management on real enterprise workflows.
- Robotics evaluations that move from controlled, near-out-of-distribution adaptation and simulations to long-duration field deployments with explicit safety and recovery metrics.
- Court developments in the publishers’ claims against OpenAI and Microsoft, and any evidence bearing on training-data use or alleged output reproduction.
Sources: S39
Sources
- Sponsored: What makes a battery a data center battery? — Data Center Dynamics · feed-summary ·
- Stulz launches new air cooling systems for Edge data center — Data Center Dynamics · feed-summary ·
- Using waste heat from data centers — Data Center Dynamics · feed-summary ·
- Ireland weighs cut-price gas for data centers in exchange for curtailment - report — Data Center Dynamics · feed-summary ·
- Thailand pauses construction on 49 data centers, as it plans new regulations — Data Center Dynamics · feed-summary ·
- Logging and Observability Guide Review Part 3 | Cloud Robotics WG Meeting 2026-10-05 — Open Robotics Discourse · feed-summary ·
- Data center development halted in Masovian Voivodeship, Poland — Data Center Dynamics · feed-summary ·
- Humain breaks ground on Neom’s Oxagon data center — Data Center Dynamics · feed-summary ·
- Greenfield and Finsbury Infrastructure to invest €360m in Milan data center — Data Center Dynamics · feed-summary ·
- Could robots help tackle loneliness? BBC’s Ann Droid raises questions about the future of care — Robohub · feed-summary ·
- Amazon Leo files for three new 6-antenna ground stations in Florida, Minnesota, and Texas — Data Center Dynamics · feed-summary ·
- FMTX: Lazy Wavefront Search for Dynamic Replanning (C++/ROS 2 framework and benchmarks) — Open Robotics Discourse · full-text ·
- N-able patches max severity N-central flaw amid ongoing attacks — BleepingComputer · feed-summary ·
- Nebulon Enterprise Simulated Threats for Phishing Research (NEST-Phish): A Synthetic Enterprise Phishing Email Dataset for Behavioral and Machine-Learning Research — arXiv Cryptography and Security · partial-text ·
- Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI — arXiv Robotics · partial-text ·
- Engineered Persuasion: Evaluating Personalized Pretexts in LLM-Generated Spear Phishing — arXiv Cryptography and Security · partial-text ·
- From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance — arXiv Artificial Intelligence · partial-text ·
- Pack It My Way: Triadic Human-Robot Collaboration for Personalized Autonomous Packing — arXiv Robotics · partial-text ·
- Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security — arXiv Artificial Intelligence · partial-text ·
- AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision — arXiv Robotics · partial-text ·
- Scalable Edge-assisted Fusion and Path Prediction for Connected Autonomous Vehicles — arXiv Robotics · partial-text ·
- Continuous Cognitive Coverage for Autonomous Robots via Event-Dependent Cognitive Treatment and Learning — arXiv Robotics · partial-text ·
- Why Is SHAP Not a Reliable Standalone Explanation Framework for Malware Detection? — arXiv Cryptography and Security · partial-text ·
- FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models — arXiv Robotics · partial-text ·
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation — arXiv Artificial Intelligence · full-text ·
- Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution — arXiv Cryptography and Security · partial-text ·
- A Removal Based Approach to Improve LLM Faithfulness at Test-Time — arXiv Artificial Intelligence · partial-text ·
- Open-Set 3D Scene Graphs for Field Robotics: An Outdoor Case Study — arXiv Robotics · partial-text ·
- Dressing in Motion: A Human Motion-Aware Diffusion Policy for Robot-Assisted Dressing — arXiv Robotics · partial-text ·
- SocioGesture: Real-Time and Adaptive Social Gesture Perception for Human-Robot Interaction — arXiv Robotics · partial-text ·
- Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets — arXiv Artificial Intelligence · partial-text ·
- Iris: Climbing to the Search Frontier — arXiv Artificial Intelligence · partial-text ·
- NASA Space Roboticist Challenge: propose an experiment for a 7-DoF arm in orbit (U.S.; registration closes Sep 23) — Open Robotics Discourse · partial-text ·
- What is the hardest part of learning ROS 2 as a beginner — Open Robotics Discourse · feed-summary ·
- Dúvida sobre organização de launch files no ROS 2 — Open Robotics Discourse · feed-summary ·
- SLAM benchmark on Livox Mid-360: public dataset + our own UAV flight data (KISS-ICP, DLIO, FAST-LIO2, GLIM) — Open Robotics Discourse · partial-text ·
- The Rogue AI Story Was Never Just A Warning Shot Or A Marketing Stunt — Forbes Innovation · feed-summary ·
- Supporting independent journalism in Ukraine — OpenAI News · feed-summary ·
- Seattle Times and Newsday sue OpenAI and Microsoft for infringement — The Verge · partial-text ·
- How Apple’s iPhone 18 Pro Will Change Smartphones Forever — Forbes Innovation · feed-summary ·
- Critical MikroTik Vulnerability - Patch Now, (Sun, Sep 6th) — SANS Internet Storm Center · feed-summary ·
- When AI Saves Time At Work, Who Gets To Keep It? — Forbes Innovation · feed-summary ·
- New CLI tool: fastdds_transport_viz — predict & measure per-topic DDS transport in ROS 2 — Open Robotics Discourse · feed-summary ·
- Eric Sivertson discusses FPGAs and robot security — The Robot Report · feed-summary ·
- Attackers conceal phishing lures using invisible Unicode characters — BleepingComputer · feed-summary ·
- Why data center cooling is now business-critical — Data Center Dynamics · feed-summary ·
- Bitcoin mining data center condemned after leaking 3 million gallons of water and forcing school closures — facility operated for years under a city stop-work order — Tom's Hardware · full-text ·
- Punching above their weight: how China’s AI giants stretch each dollar in compute race — South China Morning Post · China Tech · full-text ·
- Pressure sensors can help improve robotic gripping accuracy — The Robot Report · feed-summary ·