What are agent guardrails and how are topic boundaries and action permissions enforced?
Suggested answer
Guardrails are constraints that define what an agent can and cannot do, preventing unintended or harmful behaviour.
1. Topic boundary enforcement: Topics explicitly define scope — the Planner is instructed not to fulfil requests outside any defined topic. If a user asks about something outside all topic scopes, the agent responds that it cannot help with that request and optionally suggests escalation. Tight, well-scoped topic descriptions are the primary guardrail against scope creep.
2. Action permission checking: Each action runs in the context of the running user's Salesforce permissions. A Flow or Apex action cannot access or modify records the user doesn't have permission to — standard Salesforce sharing, FLS, and CRUD enforcement applies. The agent cannot bypass Salesforce security.
3. Einstein Trust Layer toxicity detection: Filters harmful inputs before they reach the LLM and harmful outputs before they reach the user.
4. Prompt injection protection: The Trust Layer and topic architecture provide defence against prompt injection attacks where users attempt to override agent instructions through crafted inputs.
5. Escalation thresholds: Configure the agent to always escalate for sensitive operations (e.g., processing refunds above a threshold, changing account ownership) regardless of the LLM's confidence.
Practice content for interview preparation; not an official vendor answer. Verify details against current product documentation.
Community comments (0)
No comments yet.
Sign in or create a free account to add a comment. Comments are moderated before they appear.