You're only seeing 20% of the report.
You're only seeing 20% of the report.
A practical guide to the complete agent development lifecycle.
It’s never been easier to build an AI agent. Every day, the tools and frameworks available to builders become a little more powerful, a little more intuitive to use. It’s why your social feed probably looks like a highlight reel of impressive-seeming agentic pilots and demos. But spinning up a glossy vibe-coded prototype was never the hard part. Constructing an agentic enterprise, where humans and reliable AI agents work side-by-side — that’s where the real challenge lies.
Consider these basic questions: Once an agent is deployed, who’s responsible for testing it to make sure it doesn't go off the rails when it encounters an edge case? Who evaluates performance over time? What metrics are important to track? What are the processes that meaningfully move the needle forward? The truth is, most organizations haven’t yet defined these roles and responsibilities, leading to a critical ownership gap. Without accountability, AI tooling can only go so far.
Chief AI Officer
Translating business objectives into measurable agent outcomes
Agent Builder
Implementing the agent's core functionality and ensuring technical quality
Agent Tester
Ensuring agents perform reliably across diverse scenarios before reaching users
Agent QA Lead
Identifying and diagnosing agent quality issues in pre-production
Agent Supervisor
Regularly monitoring live agent performance and identifying issues requiring intervention
Most guides focus on just building. This guide unpacks the complete agent development lifecycle (ADLC), detailing the roles and processes you’ll need to be successful post-deployment. Unlike point solutions or hyperscalers that force you to stitch together fragmented tools, Agentforce provides full lifecycle management on one unified platform. What follows is an end-to-end look at what it takes to operate enterprise-grade AI agents. We’ll explore each aspect of the ADLC, the personas accountable for each stage, the metrics they should be watching, and the actions they can take to course-correct an underperforming agent.
Diagram illustrating the complete Agent Development Lifecycle (ADLC) with its six core stages.
The planning stage is for establishing clear objectives and well-defined success criteria before building begins. This phase of the ADLC is led by:
Chief AI Officer
Often held by: CIO, Chief Digital Officer, LOB Exec
Responsible for: Translating business objectives into measurable agent outcomes
The Chief AI Officer is responsible for establishing clear objectives and well-defined success criteria before building begins. AI is a virtually limitless technology, so it’s essential to align those limitless capabilities to specific business outcomes with measurable goals that drive real impact.
Containment and deflection
Reduce the volume of cases escalated to human agents by resolving issues autonomously. When agents can handle routine inquiries from start to finish, human experts are freed to focus on complex problems that genuinely require their expertise.
Cost savings
Reduce operational expenses through automation and efficiency gains. Every interaction an agent handles successfully represents a fraction of the cost of a human-handled case, with savings that compound as the agent scales.
Revenue generation
Drive sales, upsells, and customer lifetime value through autonomous engagement. The best agents don't just answer questions — they identify opportunities, recommend relevant products, and guide customers toward higher-value solutions.
Customer satisfaction
Improve experience metrics like CSAT, NPS, and sentiment scores. When agents provide fast, accurate, 24/7 support, customers feel heard and helped, regardless of when they reach out.
Employee productivity
Free up human employees to focus on high-value, complex work. By automating repetitive tasks, agents allow your team to spend more time on work that requires empathy, creativity, and deep expertise.
The challenge is translating these objectives into concrete, measurable outcomes. That’s where a comprehensive measurement framework comes in.
Before you can measure how much ROI your agent is delivering, you need to first ensure that it’s generating good responses and that users are actually adopting it. Consider starting with a four-tiered measurement framework from the outset. Each tier answers a different fundamental question about agent performance.
Tier 4
Business value metrics
Is the agent delivering ROI?Tier 3
Engagement metrics
Are users adopting and continuing to use the agent?Tier 2
Performance metrics
Does the agent operate efficiently without degrading user experience or exploding costs?Tier 1
Efficacy metrics
Does the agent produce accurate, relevant, complete responses?
These metrics map directly to how CIOs evaluate technology partners: Efficacy proves innovation and reliability, performance validates infrastructure fit, engagement demonstrates user adoption, and business value delivers measurable ROI. We’ll explore each of these metrics in detail further down, but establishing this framework upfront helps ensure alignment later on.
Success in the planning phase requires deliberate alignment and planning.
Check out the Identify AI Use Cases unit on Trailhead for a practical framework on how to build a prioritized backlog, assess technical feasibility, and identify quick wins to jumpstart your organization’s AI transformation.
With clear success criteria established, you’re ready to build.
The build stage is for translating business objectives into working agents, implementing core functionality, and ensuring technical quality. This phase of the ADLC is led by:
Agent Builder
Often held by: Software Engineer, Platform Engineer, Technical Architect, CRM Administrator, CRM Developer
Responsible for: Implementing the agent’s core functionality and ensuring technical quality
Agent Builders translate business objectives into working agents, assigning subagents and actions, connecting data sources, configuring instructions and rules to reflect business policies, and ensuring agents can accomplish their intended use case. Remember, making something that works is only table stakes. The real goal is building an agent that scales, maintains accuracy, adheres to guardrails and doesn’t go haywire when it encounters an edge case.
Agentforce Builder
Agentforce Builder is the primary workspace for agent development. It enables users to configure the agent’s core functionality, including assigning subagents and actions, setting up data connections (e.g., knowledge bases, Data 360), defining natural language agent instructions, and implementing scripted logic for deterministic outcomes. This workspace is also where you’ll configure permission sets and mandatory compliance rules, ensuring your agent operates securely and adheres to policy.
Trust layer & security
Data is secured at every connection point, validated for quality, and governed throughout its lifecycle. Agentforce treats security as baseline infrastructure, proactively protecting your technical ecosystem from data leaks, tampering, and AI-driven threats at the source.
Subagents and actions
Subagents represent jobs to be done, such as checking an order status, resetting an account password, or answering a product question. Actions are the discrete operations the agent performs to fulfill those requests. Agent Builders can leverage standard actions that ship out of the box or create custom actions tailored to their business needs. Actions can incorporate existing Flows and Apex code, allowing organizations to fold existing processes into agentic workflows.
™ Target: All use cases from planning mapped to subagents and actions
Agent actions
Instructions are natural language rules that define the agent’s role, behavior, and guardrails. These instructions guide how the agent interprets user intent and responds to requests.
™ Target: Instructions tested and validated for clarity and effectiveness
Agentforce Script
Configured within Agentforce Builder, Agentforce Script delivers deterministic control that ensures agents act in accordance with your defined business logic. When an outcome needs to be guaranteed, such as an order return, Script ensures the agent follows predefined paths using if/then conditionals that blend LLM creativity with deterministic control.
™ Target: All compliance-critical workflows use Agentforce Script for guaranteed outcomes
Data connections
Builders connect agents to required data sources including knowledge bases, Data 360, CRM objects, and third-party APIs through configured integrations.
™ Target: All authoritative data sources connected and accessible
Permission sets
Access controls determine what data your agent can read and modify, balancing capability with security.
™ Target: Least-privilege access configured for all agent operations
Prompt Builder is a tool for creating reusable prompt templates that can be embedded across the Agentforce 360 Platform. Prompt templates are enriched with your own data and help ensure more accurate and relevant responses across multiple agents and use cases.
Hallucination risk Without proper grounding in enterprise data, agents confidently state incorrect information.
Permission gaps Agents fail to access necessary data or inadvertently expose restricted information.
Brittle actions Actions that work in testing but fail in production due to missing error handling or edge cases.
Risk aversion Overly restrictive agents create a bot-like experience that leads to poor user experience and low engagement.
But as we already established, building an agent is just the beginning. Before it interacts with real users, it needs rigorous testing.
The test stage is for systematically validating that the agent performs reliably across diverse scenarios before reaching users. This phase of the ADLC is led by:
Often held by: Quality Assurance Engineer, Software Test Engineer, DevOps Engineer
Responsible for: Ensuring agents perform reliably across diverse scenarios before reaching users
Agent Testers validate that your agent works the way you intended. They systematically probe for weaknesses, edge cases, and failure modes, catching problems in a controlled environment before real users encounter them. Testing an enterprise agent means going beyond does it work? to does it work reliably, safely, and consistently? The goal is to build confidence that your agent will behave predictably in production, while identifying any gaps that need to be patched before launch.
Testing Center
Testing Center is purpose-built for agent validation, providing AI-assisted test case generation and the ability to run tests at scale. Within Testing Center, teams can:
Accuracy
Does the agent provide correct, factually grounded responses? Are actions invoked with the right parameters?
™ Target: 95%+ correctness on test scenarios€
Scenario diversity
Can the agent handle queries across a variety of scenarios and phrasing variations? What’s the scope of what it can and can’t do?
™ Target: 90%+ pass rate across all test cases€
Error handling
When the agent encounters tool invocation errors, back-end errors, or loops, does it fail gracefully with helpful messages?
™ Target: 100% of error scenarios handled gracefully€
Performance
Do responses arrive within acceptable latency thresholds?€
™ Target: <2 seconds for 90th percentile response time
Insufficient scenario testing Testing only happy paths means edge cases surprise you in production.
Synthetic test data Tests pass with clean data but fail with messy real-world inputs.
Permission blind spots Tests run with admin access, missing permission errors real users will encounter.
Performance surprises Single-user testing doesn’t reveal latency issues under load.
Testing reveals what works and what doesn’t, but understanding why requires deeper evaluation.
The evaluate stage is for reviewing pre-production performance and diagnosing agent quality issues before deployment. This phase of the ADLC is led by:
Often held by: System Administrator, Developer, Business Analyst, Product Manager
Responsible for: Identifying and diagnosing agent quality issues in pre-production
Testers ask does it work? while Agent QA Leads ask where specifically does it fail, and why? This phase is all about surfacing specific problems that need fixing before the agent reaches users. Teams review the agent’s pre-production performance to find the most acute symptoms — subagents that consistently produce low-quality responses, queries that trigger incorrect actions, or workflows that hit latency bottlenecks. The goal isn’t just to measure quality but to pinpoint exactly what needs to be fixed and route the problem to the right person.
Testing Center
Testing Center serves as the primary environment for pre-production evaluation. Beyond generating and running tests, it enables systematic diagnosis:
Testing Center comes with a diverse range of out-of-the-box evaluations. Learn more here.
Identify negative response patterns
Surface the percentage of sessions that lead to negative outcomes (escalations, abandonments, errors) and understand which subagents correlate with those failures.
™ Target: <5% of sessions result in negative outcomes
Assess subagent-level efficacy
Evaluate which subagents consistently produce accurate, complete, relevant responses versus which ones struggle.
™ Target: All subagents meet 90%+ efficacy threshold
Diagnose latency issues
Identify subagents with high action latency that may indicate back-end loops, tool invocation errors, or inefficient action configurations.
™ Target: All subagents respond within acceptable latency thresholds
Route problems to the right persona
Determine whether issues stem from instruction clarity (send to Agent Builder), tool invocation errors (send to Developer), or knowledge gaps (send to Corpus Manager).
™ Target: 100% of identified issues assigned to appropriate owner
Symptom-focused fixes Testing only happy paths means edge cases surprise you in production.
No clear ownership Quality issues get identified but no one knows whether it’s an instruction problem, data problem, or code problem.
Evaluation happens too late Teams discover systemic issues only after deployment when fixes are more costly.
Insufficient pre-production data Limited test scenarios don’t reveal the patterns that emerge with real user diversity.
Evaluation identifies what’s broken and who should fix it. Observation monitors how it performs at scale with real users.
The observe stage is for regularly monitoring live agent performance in real time and tracking metrics to ensure the agent consistently delivers strong results. This phase of the ADLC is led by:
Often held by: Operations Lead, DevOps Engineer, Product Owner
Responsible for: Regularly monitoring live agent performance and identifying issues requiring intervention
Agent Supervisors have the critical job of monitoring production agents in real time and tracking how agents perform with actual users. Supervisors monitor the metrics established during planning (efficacy, performance, engagement, and business value) to ensure the agent consistently delivers strong results. When metrics begin trending in the wrong direction, observation helps pinpoint the cause. AI’s stochastic nature means that new errors and hallucinations can emerge at any time, so it’s important to monitor agents on an ongoing basis and label any issues.€ The goal is to know what’s happening with your agents at any moment and catch problems before they escalate.€
Monitoring metrics tells you what is happening — conversation labeling tells you why. Reviewing and categorizing interactions helps teams identify patterns and diagnose quality issues beyond what metrics alone reveal. This feedback requires dedicated ownership: assign specific roles to review transcripts and provide structured evaluation. As the feedback loop matures, many teams evolve from manual labeling to AI-assisted evaluation using custom criteria.
Agentforce Observability
Agentforce Observability provides near real-time and historical visibility into agent performance for chat and voice agents. It comprises:
Agent Analytics
Visibility into agent performance and business outcomes through pre-built dashboards for Service Agent, Employee Agent, SDR Agent, and ITSM Agent. Built on the Session Tracing Data Model (STDM), these dashboards surface metrics including deflection rates, escalation patterns, resolution efficiency, and user feedback. Analytics enables teams to measure adoption trends, identify high-performing agents, and quantify business impact across deployed use cases.
Agent Health Monitoring
Real-time detection of production issues through automated alerting. The system tracks error rates, latency, and escalation frequency against configurable thresholds. When anomalies are detected — such as error spikes or performance degradation — alerts are delivered via email, enabling teams to investigate and resolve issues before customer impact scales. Health dashboards provide historical incident data and current agent status for ongoing monitoring.
Agent Optimization
Analysis of production conversations to identify improvement opportunities. The system automatically clusters user intents daily, surfacing knowledge gaps, unresolved sessions, and unanticipated customer requests. Teams can explore individual sessions through detailed session tracing, review quality scores, and use these insights to refine agent instructions, expand knowledge coverage, and address emerging use cases based on actual customer behavior.
Observability for Voice
Unified observability across chat and voice modalities. Voice-specific capabilities include integrated media playback for reviewing full conversations and user interruption tracking to analyze conversation dynamics. This multimodal approach enables consistent monitoring, analysis, and optimization whether customers engage via voice or chat channels.
Agentforce Observability monitors agent performance across the four metric tiers established in the planning phase:
Track correctness/faithfulness, completeness, and relevancy to ensure agents maintain response quality with real user queries.
Correctness/Faithfulness: The percentage of responses that are factually accurate based on your knowledge sources.
Completeness: The percentage of queries that are fully addressed without requiring follow-up.
Relevancy: The percentage of responses that are useful to the user’s intent.
™ Target: Maintain 95%+ faithfulness, 85%+ completeness, 90%+ relevancy€
Monitor planner latency (time to decide what to do) and action latency (time to execute) to identify slowdowns before they impact user experience.
™ Target: <2 seconds for 90th percentile response time
Track active users, retention rates (percentage who return after first interaction), and session depth to understand adoption and sustained usage.
™ Target: 10%+ retention (good), 30%+ retention (excellent)
Measure containment rate (percentage resolved without escalation), time savings delivered to human agents, and cost per resolution compared to human-handled cases.
™ Target: Metrics align with business objectives from planning phase
Observation reveals when performance degrades. Iteration fixes the root causes and ensures your agents continue to improve.
The iterate stage is for validating that business objectives were met and driving ongoing improvements based on performance insights.
This phase of the ADLC is led by: Chief AI Officer + Agent Manager + AI Ops Manager
Often held by: CIO, Chief Digital Officer, LOB Exec
Responsible for: Validating that business objectives were met and driving ongoing improvements
With your agent now running reliably in production, the ADLC circles back to the Chief AI Officer, who must now validate whether the agent lived up to expectations. This determines whether the agent delivered ROI and informs future investment decisions. If the goal was 30% containment, did we reach it? If the goal was improved CSAT, did scores rise? At this phase, the Chief AI Officer is supported by an Agent Manager and AI Ops Manager to surface agent failures and rank them by priority. Which subagents underperform? Where do users drop off? What instruction changes improve efficacy? Iteration transforms agents from static deployments into living systems that get better over time.
The same stakeholders who set success criteria during planning now evaluate results:
Review whether the agent met strategic objectives: Was containment achieved? Did cost per resolution decrease? Did revenue attribution materialize?
™ Validation: Business case metrics from planning are met or exceeded
Assess whether the agent operates reliably and securely at scale: Are we maintaining uptime? Are costs predictable? Is the agent compliant?
™ Validation: Technical and operational requirements are satisfied
Evaluate whether the agent truly helps their teams: Do human employees have more time for complex work? Is user sentiment positive? Are escalations manageable?
™ Validation: Employee and customer satisfaction targets are achieved
If validation succeeds, stakeholders expand scope: Add subagents, serve more users, automate additional workflows. If validation reveals gaps, they adjust strategy, refine instructions, or revisit the business case. Either way, the cycle continues, with each iteration building on the last.
Iteration leverages insights from the observe phase to make targeted improvements:
Building an AI agent is easy. Anyone can spin up an impressive demo in an afternoon, but creating intelligent systems that reliably deliver business value, scale across the enterprise, and improve over time requires a fundamentally different approach. Without clear ownership, measurable success criteria, and continuous feedback loops, AI is little more than a shiny distraction.
The organizations that succeed in this era won’t be the ones that deploy the flashiest agentic pilots, but the ones that master the ADLC and invest in the people and processes that make it work. That’s what separates agents that thrive in production from those that fall apart under real-world complexity.