跳到主要内容
洞察
Inside

An Agent That Works for Weeks Needs More Than Memory

Salesforce's long-horizon runtime lets an Agentforce agent pursue a goal for weeks. Memory, durable execution and dynamic steering are documented. The authority questions are not.

作者 Adam Maguire Wilson约 12 分钟读完
The lighthouse on Fidra lit against a dark night sky, its lamp reflected in the sea below, standing in for an agent that must hold its course across weeks of changing conditions.
本页目录

On 11 September 2026, four days before Dreamforce, Salesforce announced a new portfolio of seven named Agentforce agents. Underneath them sits the more interesting piece of machinery: a long-horizon runtime that lets an agent "pursue goals across days and weeks instead of completing only a task or interaction". The portfolio is mostly packaging. Six of the seven agents, from Casey the help agent to Fin the customer agent Salesforce finished buying from Intercom the day before, are generally available now. The interesting one is not. Hunter, the outbound sales agent, is in pilot with general availability promised for November 2026, and Hunter is the only agent running on the new runtime.

That ordering matters, so let us be precise before anything else. The long-horizon runtime is real, announced and undocumented. Hunter is a pilot. Every other Agentforce agent you can buy today runs on the existing platform, and Salesforce has said only that more agents will move to the new runtime "over time". If you read coverage suggesting Agentforce agents now work autonomously for weeks, that is one pilot agent, not the portfolio.

We have read the announcement, the Agent Script repository and Salesforce's documentation and earnings material. We have no pilot access and have not run Hunter, so treat this as a desk read of what Salesforce has put in public, with the gaps named as gaps. Our editorial standards set out how this publication handles vendor evidence, and who we are.

Key Takeaways

  • Salesforce's long-horizon runtime rests on three documented capabilities: memory across sessions, durable execution that keeps plans running, and dynamic steering from user feedback. All three describe continuity. None of them describes authority.

  • An agent pursuing a goal for weeks will outlive the instructions, priorities and sometimes the permissions that launched it. Who can change the goal, whose steering wins when two people disagree, how you cancel cleanly, and what the audit record shows afterwards are the operative questions, and Salesforce has not yet answered them in public.

  • The inspectable part of the answer is Agent Script, the open-source behaviour language: deterministic hooks, typed state and a clean split between what the agent is and how the runtime executes it. The runtime itself remains closed Salesforce infrastructure.

  • Salesforce's own benchmark research found leading agents succeeding on roughly 58% of single-turn business tasks and 35% of multi-turn ones. A weeks-long goal is a very long multi-turn task, and no long-horizon benchmark has been published.

What actually shipped on 11 September

The announcement does two things at once, and they deserve separate treatment. The first is a go-to-market move: seven agents with first names and job descriptions, sold as configured products rather than a platform you assemble yourself. Two of them arrived by acquisition. Piper is Qualified's inbound pipeline product, which Salesforce bought in April 2026 for roughly $1.2 billion according to its quarterly filing, and Fin is the renamed Intercom, a $3.6 billion agreement in June that closed on 10 September with more than 30,000 existing customers. Calling all seven one portfolio is a merchandising decision, not a technical one; Fin still runs on its own models.

The second is the engineering. Alongside the agents, Salesforce says it has built a long-horizon runtime for Agentforce, and it frames the need with a sales example. A seller asks Hunter to rescue their at-risk deals before the end of the quarter. Hunter turns that instruction into "a measurable goal", builds a plan, works out which tasks, tools and context it needs, and operates within "guardrails that define when it can act autonomously and when seller approval is required". As work unfolds over weeks, Hunter incorporates new information and adjusts the plan "while keeping the seller in control".

Salesforce also put a number on the platform's scale: 7 billion Agentic Work Units delivered across Agentforce and Slack over two years, 3.2 billion of them in the second quarter of fiscal 2027. That figure comes from the Q2 FY2027 results on 26 August, up 97% quarter over quarter. An AWU is Salesforce's own metric, defined as one discrete task accomplished by an agent: a prompt processed, a reasoning chain completed, a tool invoked. It is a vendor-defined count of activity, unaudited, and a tool invoked badly still counts. Treat it as evidence that Agentforce runs at volume, not that the volume is valuable.

Bar chart of Agentic Work Units per quarter: 771 million in Q4 of fiscal 2026 against 3.2 billion in Q2 of fiscal 2027, a figure Salesforce says grew 97 percent quarter over quarter.

Agentic Work Units per quarter, as disclosed in Salesforce's earnings releases. The metric is defined and counted by the vendor.

Three capabilities, one sentence each

The runtime announcement names exactly three underlying capabilities, and it is worth reading the definitions closely, because each one is a single sentence and each sentence is doing specific work.

Memory "preserves context and progress across sessions, so work doesn't stop when the interaction ends". Durable execution "keeps plans running over time and allows an agent to resume or course-correct as circumstances change". Dynamic steering "adapts an agent's behavior based on an individual user's feedback and direction".

Notice what these sentences have in common. All three are about continuity: the work survives the chat window, the plan survives changing circumstances, the agent stays responsive to its human. That is genuinely hard engineering, and nobody who has run agents in production will sneer at it. Stateless request-response agents forget the deal they were working on the moment the session ends; a runtime that holds a plan for weeks, resumes it after interruption and takes correction without starting over is a real step beyond the chat-shaped products most vendors still ship.

But continuity is not authority. The three sentences tell you the agent will not forget, will not stop and will listen. They do not tell you whose goal it is once the quarter moves on, what happens when two humans steer in different directions, or how the whole thing looks in an audit six weeks later. As of 14 September 2026, those sentences are also close to the entire public record: Salesforce has published no technical documentation of the runtime's semantics, and the Agent Script repository, the obvious place to look, is explicit that the runtime is the part it does not cover. The press coverage so far, from Unite.AI to ppc.land, is working from the same announcement we are.

The goal outlives the instructions

Here is the scenario the announcement invites and then leaves hanging. A seller sets Hunter a goal in September: rescue the at-risk deals before quarter end. Between then and 30 September, the world moves. One deal closes early because the champion pushed it through. Another account gets reassigned to a different rep in a territory shuffle. The sales director decides discounting above 15% now needs her sign-off. The seller himself changes his mind about which deals matter. And, this being a CRM in a large company, the seller's own permissions might change mid-quarter as well.

A weeks-long agent lives inside all of that. So the questions that decide whether a long-horizon runtime is operable, rather than merely impressive, look like this:

Who owns the goal? The announcement describes dynamic steering as adapting to "an individual user's feedback and direction". Singular. Sales pipelines are not singular. If the seller tells Hunter to keep pushing a deal and the sales director tells it to stand down, which instruction does the plan absorb? Steering that works per user needs a documented answer to multi-principal conflict, or every enterprise pilot will discover the answer experimentally.

Can you stop it cleanly? Durable execution is defined as keeping plans running and letting the agent "resume or course-correct". Nothing in the public material describes the opposite operation: cancelling a goal, and what happens to in-flight work when you do. Does a cancelled plan halt queued outreach, or does durability mean a cancelled agent finishes what it already set in motion? Resume semantics without cancellation semantics is only half a lifecycle.

Do permissions move with the plan? Salesforce's Trust Layer documentation is clear that an agent inherits the user's access: role-based controls, field-level security and sharing rules all apply, so an agent cannot read a field its user cannot read. What is not documented is when that check happens for a goal that runs for weeks. If permissions are evaluated at goal creation, the agent is working from a snapshot that may be stale by week three. If they are re-evaluated continuously, a permission change could silently amputate part of a running plan. Both designs are defensible; they have very different failure modes. OpenAI's Codex has been working through the same problem at the smaller scale of a single turn, binding turn-scoped permission snapshots to the moment a turn starts. Weeks-long goals make the question sharper, not easier.

What does the record show afterwards? The portfolio itself demonstrates that Salesforce understands this requirement. Marshall, the supply chain agent, is sold on "deterministic execution" and "an audit record of every action". Hunter's description mentions guardrails and approvals but no audit artefact. For an agent sending outreach to real customers over weeks, an operator will want the goal-level record: the original objective, each plan revision, every steering input and who gave it, every approval requested and granted. Whether the long-horizon runtime emits that record, or leaves it to be reconstructed from session traces, is unanswered.

None of this is a reason to dismiss the runtime. It is the list an evaluator should bring to the pilot. Salesforce's own research arm, incidentally, has published the strongest available argument that these questions need answers. CRMArena-Pro, its enterprise agent benchmark from May 2025, found leading models completing around 58% of single-turn business tasks but only 35% of multi-turn ones, with what the paper calls near-zero inherent confidentiality awareness. A long-horizon goal is a multi-turn task stretched across weeks of state transitions. Salesforce has published no benchmark, error rate or intervention rate for the new runtime, so the 35% figure is the best public evidence of how agents behave once the conversation stops being short, and it comes from Salesforce itself.

What Agent Script pins down, and what it cannot

The part of this stack you can actually inspect today is Agent Script, the agent behaviour language Salesforce open-sourced under Apache 2.0 on 15 April 2026 alongside its Headless 360 platform opening. The repository is unusually honest about its boundary, and the boundary is the point.

Agent Script lets a builder define an agent as a single file: typed, mutable variables for state; before_reasoning and after_reasoning hooks that gate what the model is allowed to do next; if and else branches that make business rules deterministic rather than hopeful; explicit transition statements between subagents. In Salesforce engineering's own account of the design, the motivation was that prompts saying "do X, then Y, then Z" are fine for reasoning but cannot be trusted with "security boundaries or business-critical execution paths". Authentication, permissions and workflow sequencing get deterministic hooks; the LLM keeps the parts where adaptability helps. It is the same instinct we documented in NVIDIA's memory recipe, where the model proposes and deterministic code disposes, and it remains the healthiest pattern in enterprise agent design.

Then the README draws the line: "execution is decoupled from specification". Agent Script describes what the agent is, not how the runtime executes it, and the runtime is not open source. Scripts compile to a Salesforce-internal specification that runs on Salesforce infrastructure. The spec is even frozen to outside changes until there is a path to opening the runtime, because a spec that drifts from its only executor would split the language.

Read that boundary against the long-horizon announcement and the current state of play becomes clear. Everything Salesforce has opened, the language, the hooks, the deterministic control points, governs behaviour at the level of a step or a reasoning cycle. The announced runtime governs behaviour at the level of a goal over weeks: how plans persist, resume, get steered and, presumably, get stopped. Those goal-level semantics live exactly where the public record stops, in the closed runtime. That is not a criticism of the open-sourcing, which is real and useful. It is a map of where the remaining questions physically reside.

The pilot's own numbers deserve careful reading

One more discipline before the conclusion, because the customer figures in the announcement will be quoted back at evaluators. All six deployment statistics Salesforce published are vendor-supplied, without denominators, measurement windows or definitions of resolution. Perk's figure, that Hunter builds 60% of its sales pipeline, is the one most likely to be cited as proof of the long-horizon runtime, and it cannot be that. The same 60% appears in coverage from June 2026 describing Perk's use of the "Agentforce Prospecting Agent", three months before the runtime was announced and before the agent was called Hunter. Whatever Perk's SDR team achieved, and 60% of pipeline is a serious result if it survives scrutiny, it was achieved on the pre-announcement platform, during what we now know was Hunter's pilot lineage. It is evidence that outbound agents can produce pipeline inside Salesforce. It says nothing about how a goal held for weeks behaves.

The same care applies elsewhere in the release. Salesforce's April material named Engine's service agent Ava, handling 50% of cases; September's release names the same agent Eva at the same 50%, with no explanation. Small thing, but it tells you these figures are assembled for narrative, not audited for method. The right posture for the pilot conversation this portfolio is designed to start is the same one that applies whenever an agent asks for broader production authority: treat every one of the six numbers as a lead to verify in a pilot, not a result to plan around.

What to ask before the pilot becomes a product

Hunter goes to general availability in November 2026, and Dreamforce runs 15 to 17 September, so more detail may arrive within days of this article. When it does, or if you are in the pilot now, the questions that matter are the authority questions the announcement skipped. How is a goal cancelled, and what happens to in-flight actions? When two humans steer differently, whose direction wins, and is that precedence configurable or hard-coded? Are CRM permissions re-evaluated as a plan runs, or snapshotted at goal creation? And does the runtime produce a goal-level audit record, the plan versions, steering events and approval decisions, the way Marshall's listing promises an audit record of every action?

Salesforce has built something genuinely new for its platform here, and the honesty of the Agent Script repository about what is open and what is not suggests the company knows exactly which layer it has yet to explain. An agent that works for weeks needs more than memory. It needs an answerable chain of authority that lasts as long as the goal does. We will find out this autumn whether the long-horizon runtime has one, or whether that, too, is still in pilot.

继续阅读

Agent Field Notes

获取下一期。

为需要真正运行这些系统的人,解释 Agent Harness、运行时、安全与治理。

是否正面临类似的决策?

我们为面临重大智能体系统决策的团队提供架构审查、治理评估和版本固定的框架评估。

关于作者

Adam Maguire Wilson

创始人,AI 智能体系统独立顾问。

adam.mw