This website uses cookies

Read our Privacy policy and Terms of use for more information.

A reminder that italicized text is my own authorship; otherwise, it’s analysis from AI models.

This week:

  • Slowdown? What slowdown?

  • Who owns the customer experience when that customer isn’t human?

  • Datacenter builds are getting more questions than answers.

AI This Week — Week of September 24, 2026

What moved in AI this week — plain English, weekly arc

The Big Story This Week

One week after the slowdown call, what actually changed?

Last week’s issue examined who gets to set the pace of AI development. It reported calls for restraint, not an agreed pause. This week provides an initial test of what those calls mean in practice: an embedded evaluator was named, questions about independence persisted, and a proposed testing agreement reportedly failed to materialize.

  • Anthropic named Accenture’s Faculty unit as its first embedded evaluator. Neatprompts reported that the arrangement gives evaluators access to red-team models and safeguards, but also noted that Accenture is a major commercial partner whose employees and developers are being trained on Claude. The newsletter therefore treated independence as unresolved. (Neatprompts, September 21, 2026)

  • A proposed reciprocal testing arrangement between OpenAI and Anthropic reportedly did not close. AI Daily Brief described the talks as part of a broader negotiation over whether rival labs, outside evaluators, or government rules should prevent companies from grading their own work. (AI Daily Brief, September 22, 2026)

  • Governance coverage later in the week returned to the same conflict. AI Governance, Ethics and Leadership paired the Anthropic–Accenture arrangement with emerging automated-compliance systems, while questioning whether visible evaluation is the same as independent evaluation. (AI Governance, Ethics and Leadership, September 24, 2026)

  • Product competition continued alongside the governance debate. The Signal reported Claude/Cowork consolidation and expanded Salesforce integration, while Lenny’s Newsletter reviewed Meta’s Muse consumer agent. These are examples of continued product activity; they do not establish whether frontier training slowed or whether any company breached a specific commitment. (The Signal, September 20, 2026; Lenny’s Newsletter, September 21, 2026)

The clearest follow-up is the gap between support for oversight and agreement on how it should work. Naming an evaluator makes one proposal more concrete, but access alone does not settle independence. Reported difficulty agreeing on reciprocal testing shows that coordination remains unresolved.

There is no established pause here whose duration we can measure. The week’s evidence instead shows attempts to build oversight while companies continue competing for users and workflows.

Z’s Take

After much hand-waving and letter-writing, not much was actually done in the way of slowing down or really governance. Instead, what we got this week were more updates to ChatGPT and Claude, and what seems to be a move towards having some kind of oversight internal to Anthropic.

People are naturally asking, does having Accenture sitting in desks at Anthropic count as actually having public oversight? In the end, the calls for slowdowns look a lot like marketing moves ahead of IPOs to try to look responsible and to react in a way that would soften the public’s reactions to all the reports of agents hacking Hugging Face.

What Built Momentum

Signals that became meaningfully stronger this week through wider source coverage, a new operational manifestation, or movement toward a durable long-term trend.

The agent is becoming the interface—and platforms are deciding whether to let it in

For most of the agent era, the question was whether an AI could complete a task. This week, the more consequential question was where that task begins. Several products put the agent between the user and the software, store, or service underneath it. Then one large platform pushed back.

  • Anthropic moved the starting point from the file to the conversation. The Signal reported that Claude and Cowork are merging into one app, with documents and slides created inside the same conversation. Salesforce in Claude adds 37 prebuilt sales skills, and Anthropic said 7,000 Salesforce sellers were already using it. Those adoption figures are company claims reported by The Signal. (The Signal, September 20, 2026)

  • Meta’s Muse showed what a consumer version of that interface can look like. In Lenny’s Newsletter, Muse managed a family calendar, created a morning newsletter, tracked goals, requested permission when sensitive data became relevant, and exposed an activity feed of its tool calls and steps. The same hands-on review found its browser use unreliable, so this was evidence of a product direction—not proof that the agent worked consistently. (Lenny’s Newsletter, September 21, 2026)

  • Amazon reportedly blocked Muse from shopping on its site. AI Daily Brief framed the dispute as a fight over who owns the customer relationship when an agent sits between a person and an online platform. The newsletter also reported that Muse had passed ChatGPT to reach No. 1 in the U.S. App Store; that ranking and the examples of purchases made through Muse were reported claims, not independently checked in this analysis. (AI Daily Brief, September 22, 2026)

The new development is not simply that agents can use software. It is that the agent may become the place where people begin work and commerce—while the services underneath it decide whether that intermediary helps their business or threatens it.

Z’s Take

There are two reasons everyone in Insights should be paying attention to this one.

  1. First, for insights professionals particularly, when an agent does analysis and creates a report with made up quotes and references to data that don't exist, who is ultimately responsible for that report? If synthetic data is used to do research, and someone writes a report and adds it to the organization’s research database, who is responsible for making sure that employees are able to easily tell which reports were written based on synthetic data versus reports written based on first-party, human-based data?

  2. The second reason Insights should care about this is having agents doing shopping on behalf of customers means that there are going to be questions about the change in the shopper experience. Not only that, but there's going to be a change in the ways that companies selling products are going to try to get customers' attention. Who or what are companies now trying to grab attention from? Is it people or is it AI agents?

    We've already seen the impact that AI can have on the way people interact with companies via the way web traffic was impacted by search results via AI instead of search engines like Google. Now, instead of just having to worry about SEO, marketing teams also worry about AEO or making their websites visible to AI tools so that their sites surface when someone asks a relevant question via AI. Insights pros like Yogesh Chavda have been talking about the shift in shopping for at least 6 months now; if companies haven’t been paying attention, now might be the time to take notice and strategize on what to do if AI starts shopping at your online store instead of a person.

Data-center growth is becoming a local accountability problem

Infrastructure cost has been a long-term signal. What changed this week was the concentration of evidence about who approves AI infrastructure, who bears its local effects, and what operators have to disclose.

  • Slow AI examined proposed Scottish projects at neighborhood scale. It described a proposed 200-megawatt campus on green-belt land and a separate 600-megawatt development. The newsletter said roughly 100 more sites were in planning or construction across the UK, while stressing that developer water and jobs claims were incomplete or unaudited. (Slow AI, September 18, 2026)

  • The same article reported governments struggling to count and govern the build-out. Slow AI said Thailand paused 166 proposed or active projects while assembling information across 16 agencies, and that Scotland had introduced a requirement for councils to notify ministers of validated data-center applications. These are the newsletter’s reported figures and policy descriptions. (Slow AI, September 18, 2026)

  • Nita Farahany connected the expansion to household-level power, noise, water, tax, and permitting questions. She reported that the PJM grid’s 2027–28 capacity auction came up 6,623 megawatts short of its safety margin and described both local complaints and local tax benefits. The examples show why infrastructure planning is becoming a distribution question, but they do not establish that data centers alone caused every grid or community outcome described. (Nita Farahany, September 21, 2026)

The signal has moved beyond “AI needs enormous amounts of compute.” The growing issue is whether communities can see the resource demands, negotiate benefits, and influence the decisions before a project is built.

Z’s Take

Here's another thing that I think insights professionals should be looking at. Datacenter builds are impacting the environment and the towns nearby. We're starting to hear more reports of how data centers are impacting the people who live near them, such as increased heat and increased noise in nearby towns. These data centers are having broader impacts than just electricity and water supply, but who's tracking them?

Evaluation is moving inside the workflow

Evaluation has appeared throughout the full corpus. This week it advanced from periodic model testing toward continuous operating control.

  • Warp described scoring every agent run, not just testing a model before deployment. Its software factory uses an AI judge across multiple dimensions, collects roughly 20–25 similar failures before proposing changes, and replays real work to compare model configurations. The company said its agent averaged 35 minutes from kickoff to pull request while the first human review arrived 3.5 hours later. These are company examples reported by Lenny’s Newsletter. (Lenny’s Newsletter, September 21, 2026)

  • Lenny’s Newsletter also described error discovery as the evaluation equivalent of product discovery. Reviewing real user traces can change a team’s definition of a good result—a problem called “criteria drift.” The article argued that production failures should become repeatable tests before later changes ship. (Lenny’s Newsletter, September 22, 2026)

  • AI Maker translated the same idea into everyday agent work. It recommended defining what good looks like, requiring evidence behind a verdict, and checking whether the human agrees with the evaluation rather than accepting a yes-or-no score as objective. (AI Maker, September 24, 2026)

The momentum is not another benchmark score. It is evaluation becoming part of the loop that decides whether work is complete, what failed, and what the system should change next.

Z’s Take

So, evaluations have previously been used to understand how any particular generative AI model performs for a given set of tasks within a given industry. The majority of the evaluations that I have seen have been for product engineering, product management, or software development.

I have yet to see the same rigor designed for insights, marketing, fintech, pharma, or any other of a host of industries which might be a sign of where these tools are being used most.

What Kept Showing Up

Long-term signals that continued this week without enough change in trajectory to make them momentum stories.

Human judgment remains the verification layer — Long-term sustained

  • Human+AI argued that critical thinking means examining evidence, assumptions, and alternatives rather than accepting the first plausible answer an AI produces. (Human+AI, September 22, 2026)

  • Slow AI described an automated moderation system that correctly detected threatening imagery but drew the wrong conclusion about a video criticizing that imagery. The successful appeal showed that an accurate detection can still produce a bad decision when context is lost. (Slow AI, September 23, 2026)

  • AI Maker’s evaluation workflow still ends with a person deciding whether the system’s standard and evidence are persuasive. Automating the check changes where judgment occurs; it does not remove responsibility for the standard. (AI Maker, September 24, 2026)

Using multiple models is an operating strategy — Long-term sustained

  • Every’s same-week comparisons of Claude Opus 5.5, GPT-6 Sol, and other models focused on fit by task rather than naming one universal winner. (Every, September 22, 2026)

  • AI Daily Brief made the same point in its comparison of Opus 5.5, GPT-6 Sol, and GPT-6 Luna: the useful question was where each model belongs in a working rotation. (AI Daily Brief, September 23, 2026)

The recurring pattern is model choice as routing: quality, speed, price, context, and workflow fit determine which system gets the job.

Agentic work keeps expanding beyond the prompt — Long-term sustained

  • The Signal described Claude projects splitting a goal across parallel threads with shared memory and connecting the conversation to Salesforce records and skills. (The Signal, September 20, 2026)

  • Lenny’s Newsletter described Warp’s factory moving a request from Slack through triage, ticket creation, implementation, testing, pull-request creation, and review. (Lenny’s Newsletter, September 21, 2026)

The unit of design is increasingly the surrounding workflow—tools, permissions, memory, checks, and handoffs—not the prompt alone.

What to Watch

Earlier-stage signals with enough evidence to follow, but not enough duration, breadth, or trajectory for a durable classification.

Workflow-specific models are gaining stronger examples

  • Figure said its Helix 2.5 robots completed whole household tasks in 56% of trials, compared with 9% without the company’s human-video training. The tests were company-run and not independently verified. (The Signal, September 20, 2026)

  • OpenAI launched Astra for Law, which searches more than 230 million U.S. legal sources. The Signal reported OpenAI’s own result of 54% on 200 legal-research questions, compared with 38.7% for the same model using web search; the result was not independently verified. (The Signal, September 20, 2026)

These examples strengthen the case for models adapted to particular domains or proprietary data. But because both examples came through one newsletter on one day, the signal remains on Watch rather than moving into Built Momentum.

Z’s Take

I'm going to be a bit contrarian here, but are we really celebrating the fact that a tool got a 54% on a test? Or that a robot was able to complete 56% of trials doing household chores? I'd hate to be the person who bought a robot who only succeeds in doing what I want it to barely over half the time. I already yell at the Google devices that we have in the house for only managing to play a song I ask for roughly half the time. Maybe I should ask myself why I keep using them for that purpose…

Open models may be moving from alternative to default infrastructure

  • The Signal reported that open models accounted for 78% of token volume on Vercel’s AI Gateway on September 19, compared with 22% for closed models. It also relayed estimates that DeepSeek V4.1 Flash was 45–75 times cheaper to run than GPT-6 Astra depending on token direction. The usage figure and cost estimates came from cited third parties and were not independently audited here. (The Signal, September 23, 2026)

  • The same analysis argued that proprietary systems still retain advantages in product integration, brand, and frontier capability. That makes this a competition over which layer captures value, not evidence that closed models have become irrelevant. (The Signal, September 23, 2026)

Open-model use has appeared before; the new usage figure is notable enough to track, but one gateway and one newsletter are not a market-wide trend.

Also Worth Watching

  • Industrial robotics is being measured in risk removed, not only labor saved. Neatprompts reported Chevron’s claim that robotics had removed 143,000 hours of at-risk work since 2024. This is a company-reported operational figure, not an independent safety study. (Neatprompts, September 20, 2026)

  • AI-native organization design is producing new management roles. Elena Verna described Lovable’s “agent parents,” who manage portfolios of agents and use AI to coordinate work across a company with deliberately few managers. The article was also distributed by Lenny’s Newsletter, so the two appearances are one underlying source, not independent corroboration. (Elena’s Growth Scoop, September 24, 2026)

  • Automated decisions can fail even when the underlying detection is accurate. Slow AI’s moderation example is a reminder that evidence about what a system noticed is not the same as evidence that its decision was justified. (Slow AI, September 23, 2026)

This newsletter covers Friday, September 18–Thursday, September 24, 2026. Sources reviewed include Slow AI, Neatprompts, The Signal, Unpromptable, Nita Farahany, Every, AI Daily Brief, The Voice of User, Lenny’s Newsletter, AI Maker, AI Governance, Ethics and Leadership, Human+AI, AI Leaderboard, Prompt-Led Product, On New Terms, Elena’s Growth Scoop, and Last Week in AI.

Recommended for you

View all
caret-right