What If You Hired an AI Agent Like an Employee?

AI is now sold as job roles, not features — here's how a small team should evaluate, trial, and fire one.

On September 11, 2026, Salesforce shipped a batch of Agentforce agents with job titles instead of feature names: Casey for customer support across voice, SMS, WhatsApp and web chat; Paige for internal IT and HR requests; Carter for shoppers, including product comparison and in-chat checkout; Marshall for supply chain; Piper for qualifying inbound leads; Fin for customer service; and Hunter for outbound sales. Most are generally available, while Hunter is in pilot with general availability announced for November.

The names are not the story. AI used to be sold as features — summarize, draft, chat. It is now being sold as positions on an org chart. That turns a software comparison into a hiring decision, and the smaller your team, the faster you feel it.

What actually changes for a two-person team?

The evaluation question flips. Instead of "what can this tool do," you ask "how much of this job does it finish end to end?" That favors small teams, because the people with no time to assemble features benefit most from something already packaged as a job.

Take return requests in an online store. The old path was: add a chatbot tool, load the return policy, wire up order lookup, then write escalation rules. A job-shaped agent arrives with FAQs, returns, account handling and human handoff already wired together. Salesforce connected these agents to Customer 360 for exactly that reason — they run on the customer data and processes a company already has. The flip side is blunt: if your own policies and records are messy, a job-ready agent flounders like an unbriefed new hire.

What should you decide before you buy?

Write a one-page job description before you compare vendors. Borrow the sequence you'd use for a human hire. These five items alone make both selection and cancellation far easier.

  1. Three lines of duties: "look up orders, accept returns, explain shipping delays." Anything not listed is out of scope.
  2. Authority limits: read-only, or can it issue refunds? Put a currency threshold in writing and require human approval above it.
  3. Handover material: policy docs, past tickets, product details. If you wouldn't hand it to a new hire, think twice before handing it to an agent.
  4. A probation period and pass mark: two weeks, 30 cases a day, and count the human-rework rate and escalation rate. Count it, don't feel it.
  5. Firing conditions: which metric failing turns it off, and who absorbs the work when it does.

What's risky about agents that run for weeks?

The genuinely new part of the announcement is the long-horizon runtime behind Hunter — an agent that pursues a goal over weeks rather than ending with one conversation. For a small team that's both leverage and exposure: the longer it runs unattended, the longer mistakes and costs accumulate quietly.

  • Checkpoints: once a week, a human reads and approves "here's what I did and what's next."
  • Anything a customer sees: keep send-approval on for the first two weeks. A bad outbound email can't be recalled.
  • Billing units: Salesforce defines a separate unit for counting a single task an agent completes. Before signing, ask what exactly counts as one unit — per conversation or per resolved case changes the invoice completely.
  • A kill switch: confirm that disabling one account stops every automated send.

One last angle: this trend is also a selling opportunity. Big-platform job agents are broad and average. Roles that need industry-specific phrasing and messy exception handling — dental intake, tutoring-center absence notices, contractor quote replies — are still wide open. Your task this week is simple: pick your single most repetitive job, write its description, run a two-week probation. Only then does "which AI should we buy" become a real question.

FAQ

How is a job-ready agent different from a chatbot?
A chatbot gives you a conversation layer and leaves the workflow design to you. A job-ready agent arrives with the procedures and escalation rules for one role already built in. It still needs your policy documents and customer records to be in order before it performs.
Is this worth it for a solo founder?
You don't need to staff every role — start with the single highest-volume, rule-based task you have. Returns intake or booking confirmations show measurable results within a two-week trial. Work where the rules change case by case is still faster with a human.
Are long-running agents safe to deploy?
Capability matters less than supervision structure. Weekly plan approvals, review before anything goes to a customer, and an instant kill switch remove most of the risk. Confirm the billing unit up front so a multi-week run doesn't quietly inflate your invoice.