Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

Agent TestingAI TestingAI

Agentforce Testing Center: Scope, Cost, and Sandbox Constraints

Agentforce Testing Center explained: what it tests, documented limits, Flex Credit cost, sandbox constraints, and where end-to-end agent testing fills the gaps.

Published on:

Your Agentforce agent passes every case in Testing Center. Green across subagent, action, and response. Then it goes live on the website chat, and the first real customer asks in Spanish, switches topics halfway through, and gets a case number for a case that was never created.

Testing Center did its job. It answers a narrower question than most teams assume: did the agent make the right decision inside Salesforce? End-to-end agent testing asks whether the whole thing worked for the customer, through the real channel, all the way to the record.

This guide covers Testing Center's scope, cost, and sandbox constraints, using the limits and credit charges Salesforce documents but rarely puts in one place. Where end-to-end coverage is needed, it shows how TestMu AI tests Agentforce agents through the conversation itself.

Key Takeaways

Agentforce Testing Center checks whether an Agentforce agent picked the right subagent, called the right actions, and gave an acceptable response, inside a Salesforce sandbox. End-to-end agent testing drives the same agent through the customer channel across full conversations and verifies the outcome. Most production agents need both before go-live.

  • Testing Center scope: Agentforce Testing Center grades three things per test case: the expected subagent, the expected actions, and the expected response, which an LLM judge scores from 0 to 5 with 3 or higher counted as a pass.
  • Credit cost: Running Agentforce Testing Center tests consumes requests and credits the same way customer usage does, even in a sandbox, where the Flex Credits rate card charges 16 credits per standard action instead of 20.
  • Published limits: Salesforce documents two different Agentforce Testing Center limits: 1,000 test cases per test and 10 jobs per 10 hours on Trailhead, and 500 cases per job and 10 jobs per hour in Salesforce Help.
  • Sandbox choice: Developer and Developer Pro sandboxes copy metadata only, so an Agentforce agent grounded in records needs a Partial Copy or Full sandbox, or seeded data, for Testing Center results to mean anything.
  • Language coverage: Agentforce Testing Center supports English only as of May 2026, so non-English customer conversations need a separate end-to-end test to be covered at all.
  • End-to-end coverage: End-to-end agent testing adds the real channel, multi-turn conversations, varied personas and languages, and checks on what the agent actually changed, which Testing Center does not verify.
  • TestMu AI Agentforce testing: TestMu AI connects to an Agentforce agent as an HTTPS endpoint and runs generated multi-turn scenarios with 15+ specialized AI evaluators, including scenarios in languages other than English.

What Is Agentforce Testing Center?

Agentforce Testing Center is Salesforce's built-in tool for batch-testing Agentforce agents before they reach customers. Each test case sends an utterance to the agent and checks what the agent decided. Salesforce's Trailhead unit on testing criteria defines the CSV template with four columns:

  • Utterance - what the user says, such as a request for open opportunities on an account.
  • Expected Subagent - the API name (not the label) of the subagent that should handle it. Older docs call this the topic.
  • Expected Actions - a bracketed list of single-quoted action names.
  • Expected Response - a description of what the reply should cover.

Only the utterance plus one other column is required, and empty values are treated as failures, so leave a column out entirely instead of blank. You can also let AI generate cases from the agent's subagents and actions; the default is 20. Users need both the Manage Agentforce Grids and Manage Agentforce Testing permissions in the new Testing Center.

What Is End-to-End Agent Testing for Agentforce?

End-to-end agent testing exercises the agent the way a customer does: through the deployed channel, over a full conversation, and all the way to the outcome. The general approach is covered in end-to-end agent testing; for Agentforce, the path has these stages:

  • The customer reaches the agent through a channel, such as web chat, in-app messaging, or voice.
  • The agent reasons over several turns, with context carried from earlier messages.
  • The agent calls actions such as Flows, Apex, or prompt templates.
  • Those actions change records or reach external systems.
  • The conversation ends in a resolution, a handoff to a human, or a follow-up.

Testing Center covers stages 2 and 3 in depth. Stages 1, 4, and 5 are where end-to-end testing earns its place. Salesforce's April 2026 Testing Center update narrowed the gap on stage 2 by adding conversation-level testing and personas, which the comparison below reflects.

How Do Agentforce Testing Center and End-to-End Testing Compare?

Testing Center grades the agent's decisions inside the org, and end-to-end testing grades the customer's experience and the result. Each Testing Center cell below comes from Salesforce's own docs:

DimensionAgentforce Testing CenterEnd-to-end agent testing
Entry pointUtterances sent to the agent inside Salesforce.The deployed channel or endpoint customers actually use.
Conversation depthSingle utterances, plus conversation-level tests since April 2026.Full multi-turn conversations as the default unit.
What is gradedSubagent, actions, and response; custom evaluations optional.Response quality per turn, plus the outcome in the org and downstream.
LanguagesEnglish only, multilingual planned.Whatever languages your customers use.
ScaleCapped per job and per time window (see limits below).Set by the testing platform you use.
EnvironmentSandbox; production not recommended.A sandbox or staging channel; production only for read-only checks.
Agentforce creditsConsumed the same as customer usage.Also consumed, because the agent runs for real.

The last row matters for budgeting. Any test that invokes the agent triggers Agentforce actions, so end-to-end testing doesn't avoid Salesforce credit usage; it buys coverage Testing Center can't reach. The broader distinction between grading replies and testing agents is covered in LLM evaluation vs agent testing.

What Are the Limits of Agentforce Testing Center?

Salesforce's own sources give two different sets of numbers, and it's worth planning around the stricter one:

  • Trailhead - the Trust Your Agents unit says you can run up to 10 test jobs in a 10-hour time frame and have up to 1,000 test cases per test.
  • Salesforce Help - the Learn About Agentforce Testing Center article (May 2026) lists a maximum of 500 test cases per job and 10 jobs per hour, with a recommended batch of 20 to 30 cases per evaluation.

The same Help article adds constraints that shape what you can test at all:

  • Run time - about 5 seconds per test case on average, so a 500-case job takes roughly 42 minutes.
  • Response grading - an LLM judge scores responses from 0 to 5, and a 3 or higher passes.
  • Language - English only, with multilingual support planned.
  • Agent types - Service Agent, SDR Agent, Employee Agent, and the Default Agent are supported.

Before you size a suite, open Considerations for Testing Center in Salesforce Help for your release. In practice, I'd keep regression jobs well under 500 cases and split by subagent, which also makes failures easier to read.

Does Agentforce Testing Center Consume Flex Credits?

Yes. The Trust Your Agents unit is explicit: running tests, manually or automatically, consumes requests and credits the same as your customers using the agent, and that holds even in a sandbox. There's no license fee for Testing Center itself; the cost is in the agent activity it triggers.

For contracts on Flex Credits, the Flex Credits Rate Card (June 17, 2026) sets these multipliers per action:

Usage typeProductionSandbox
Standard or custom Agentforce action20 credits16 credits
Agentforce voice action30 credits24 credits

The rate card applies the sandbox multiplier to all pre-production environments, including scratch orgs, and lists no exemption for testing. As a worked example, a 1,000-case test where each case triggers two standard actions is 2,000 actions, or 32,000 Flex Credits in a sandbox. Contracts on older Einstein Requests show up the same way: Salesforce Help says to track consumption in Digital Wallet.

Run a 10-case job first, then check Digital Wallet before you schedule anything large. That single number tells you what your contract actually charges per test case.

Note

Note: TestMu AI runs Agentforce conversations as full multi-turn scenarios, in the languages your customers use, with 15+ AI evaluators scoring each one. Start testing for free.

Which Sandbox Should You Run Agentforce Tests In?

A sandbox, always. Trailhead warns that testing agents can modify CRM data, and Salesforce Help marks production enablement as not recommended. The harder question is which sandbox, because the sandbox type decides what data the agent can ground on. Salesforce's sandbox licenses and storage limits table:

Sandbox typeWhat's copiedRefresh intervalFit for agent testing
DeveloperMetadata only (200 MB data storage)1 dayRouting and instruction tests with seeded records.
Developer ProMetadata only (1 GB data storage)1 daySame, with room for a larger seeded dataset.
Partial CopyMetadata and sample data (5 GB)5 daysGrounded agents, if the template includes the objects they read.
FullMetadata and all data29 daysPre-release and end-to-end runs closest to production.

The trap is the Developer sandbox. It has your agent's metadata and none of the records it answers from, so an order-status agent fails every grounded test for a reason that has nothing to do with the agent. Either seed the objects the agent reads, or run grounded tests in a Partial Copy or Full sandbox. Each test also leaves real records behind, so reset or refresh before you compare one run with the next.

What Does End-to-End Testing Catch That Testing Center Misses?

Testing Center tells you the agent chose well. These are the failures that only show up when you test the whole path:

  • Channel configuration - the chat deployment, pre-chat fields, or authentication context your channel passes in, none of which exist in a Testing Center utterance.
  • Non-English customers - Testing Center's English-only support leaves these conversations untested.
  • Long conversations - topic switches, corrections, and context carried across many turns.
  • The outcome - the case, refund, or email the reply promised, and whether it exists with the right values.
  • Handoffs - whether an escalation actually reached a human queue.

TestMu AI Agentforce testing covers the conversation half of that list. It connects to the agent as an HTTPS endpoint (or through a secure proxy for private networks), generates 60 to 100+ scenarios per workflow from the documents you upload, up to 300 per generation run, and runs each one as a full multi-turn conversation with 15+ specialized AI evaluators. Scenario generation and evaluation are multilingual, with the language chosen at generation time.

TestMu AI Agent Testing scenario for graceful fallback on an unsupported intent, showing the user input, the expected response, and three validation points mapped to evaluator agents

The screenshot above is a generated error-handling scenario from the TestMu AI platform: the user asks something outside the agent's scope, and three validation points check that the agent recognizes the unsupported intent, states its scope politely, and offers supported alternatives. In Testing Center, the same check is a CSV row you write and maintain yourself.

For the outcome half, check the records after each end-to-end run, or verify the effect directly. Agentforce regression testing covers how to wire both into a pipeline.

Austin Siewert

Austin Siewert

Co-Founder, Steadfast Systems

Discovered @TestMu AI yesterday. Best browser testing tool I've found for my use case. Great pricing model for the limited testing I do 👏

2M+ Devs and QAs rely on TestMu AI

Deliver immersive digital experiences with Next-Generation Mobile Apps and Cross Browser Testing Cloud

When Should You Use Each Approach?

Use both, at different moments. Given Testing Center's 5-second average per case and its per-job cap, it suits fast checks after each change; end-to-end runs suit release gates.

SituationStart withWhy
Edited a subagent instructionTesting CenterRouting and action checks answer the question in minutes.
Changed a Flow or Apex actionBothTesting Center confirms the call; end-to-end confirms the record.
Launching on a new channelEnd-to-endChannel setup is invisible to Testing Center.
Serving non-English customersEnd-to-endTesting Center is English only.
Pre-release sign-offBothDecisions and outcomes both need a pass before go-live.

For the Salesforce side of the stack, from Apex tests to UI automation, the Salesforce testing guide covers what surrounds the agent, and the AI agent testing guide covers the evaluation methods behind both approaches.

How to Start Testing Agentforce End to End

Start in a Partial Copy or Full sandbox with a 10-case Testing Center job, and check Digital Wallet so you know the credit cost per case. Keep that suite for routing and action checks after every change.

Then point TestMu AI at the same agent through its channel endpoint. The chat agent API integration docs show how to connect a public endpoint or a private one through the secure proxy, so your first multi-turn, multilingual run covers what Testing Center can't.

Author

...

Samyak Goyal

Blogs: 28

  • Linkedin

Samyak Goyal is a Senior Member of Technical Staff at TestMu AI engineering Kane CLI, the command-line tool that runs browser automation from the terminal, where a flow described in natural language executes in a real Chrome browser and returns pass or fail with shareable proof. He is a backend engineer with 4+ years of experience, previously an SDE at Innovaccer, where he built APIs, introduced Kafka, and cut deployment from weeks to hours. Samyak also builds multi-agent systems, skill-orchestration frameworks, and a personal copilot that indexes 200+ microservice repositories.

Reviewer

...

Vipul Verma

Reviewer

  • Linkedin

Vipul Verma is Group Senior Vice President of Engineering at TestMu AI (formerly LambdaTest), where he heads the entire engineering organization that builds KaneAI, HyperExecute, and the broader testing cloud. He brings 15+ years architecting, securing, and scaling large enterprise applications across multiple sites. Before TestMu AI he was India Head at LogicHub, where he built the India R&D site from the first employee to a 30-plus engineering team, and Principal Software Engineer at Sumo Logic, where he was the first engineer in the India office and shipped search-performance and pricing-model initiatives. Earlier he worked on trading platforms at Portware and D. E. Shaw. Vipul holds a B.Tech in Computer Science from IIT Kharagpur.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Agentforce Testing Center FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests