Last Updated on 22 Sept 2026
Your System Prompt Is Not a Vault: Keeping Secrets, Authorization, and Security Policy Outside the LLM
Share in

Introduction
System prompts help define how an LLM application should behave. They can describe the assistant's role, response style, limits, available tools, and business context. Because users do not normally see these instructions, teams may begin treating the prompt as a hidden policy file, a secret store, or an authorization boundary.
OWASP identifies system prompt leakage as LLM07:2025. A prompt may be exposed through direct requests, prompt injection, model behavior, debugging, logs, connected tools, or application mistakes. More importantly, OWASP advises organizations not to rely on system prompts for strict behavior control when enforceable systems outside the model can perform that job.
The safest architecture assumes that the prompt may eventually become visible. If disclosure reveals a writing style or public workflow, the impact should be limited. If disclosure exposes credentials, private keys, internal access paths, customer secrets, or the only rule preventing a financial action, the architecture has placed real security inside a probabilistic and potentially observable component.
What belongs in a system prompt
A system prompt is appropriate for role guidance, response tone, task boundaries, formatting preferences, clarification behavior, and instructions about when to ask for help. These elements improve the user experience and reduce ambiguity. They influence the model, but the application should remain safe if the model fails to follow them.
The prompt can also describe available tools and high level policies, provided the real enforcement lives elsewhere. For example, it may tell an assistant to request approval before a refund. The workflow engine must still prevent the refund until approval exists. The instruction explains expected behavior, while deterministic code enforces the requirement.
Teams should write prompts with disclosure in mind. Internal wording can remain confidential for commercial reasons, but confidentiality should not be the control that protects users or systems. A useful review question is whether publishing the prompt would create unauthorized access or expose a credential. If the answer is yes, the sensitive material should be moved.
What should never be stored in a system prompt
Secrets do not belong in prompts. API keys, passwords, tokens, private keys, database credentials, recovery codes, and signing material should be held in a secrets manager and supplied only to the component that needs them. The model should receive a narrow tool capability rather than the credential used to perform the capability.
Sensitive customer data should also be excluded unless it is necessary for the immediate task and permitted for the current user. A prompt shared across sessions should never contain one customer's information. Dynamic context needs access checks, minimization, retention rules, and separation from the static instructions that configure application behavior.
Security controls should not exist only as natural language. A line telling the model never to reveal data, never to exceed a limit, or never to call an unsafe tool is not equivalent to data authorization, a budget, or a tool permission. Prompt instructions can support defense, but they cannot replace enforceable boundaries.

Moving authorization and policy outside the model
Authorization should be evaluated by the application for every protected resource and action. The decision can use the authenticated identity, role, tenant, ownership, requested action, current state, and risk signals. The model may recommend an action, but it should not decide whether its own credentials permit that action.
Tool access should use narrow interfaces and separate credentials. A refund tool can expose an operation with an amount ceiling instead of general payment administration. A search tool can return only records already filtered for the user. The application can bind identifiers from the trusted session and ignore model generated attempts to substitute another account or tenant.
Policy engines are useful when rules change or need auditability. They can record which version approved an action and provide a consistent decision across models. Human review can remain part of the policy for sensitive exceptions. This design makes the security result explainable without depending on whether a model interpreted prose exactly as intended.

Protecting tools, data, and execution
Each tool should have the minimum network, data, and action scope required for its task. Read functions should be separate from write functions. Reversible actions should be preferred when possible. High impact actions should require stronger authentication, transaction limits, or approval, even when the model response appears confident.
Input and output validation protect the boundary on both sides. Arguments created by the model need schema checks, allow lists, ranges, ownership verification, and context specific encoding. Tool responses should be minimized before returning to the model so one successful call does not expose more information than the task requires.
Logging should capture tool selection, validated arguments, authorization, risk signals, approval, result, and policy version. The log should not reproduce secrets or unnecessary sensitive data. These records help teams investigate whether a failure came from prompt manipulation, authorization, tool design, session compromise, or downstream execution.
Classifying prompt content by consequence
Prompt reviews should classify instructions by what happens if they are disclosed or ignored. Style, tone, and formatting guidance usually create limited security impact. Internal workflow descriptions may reveal useful reconnaissance. Credentials, customer data, authorization logic, transaction limits, and private network details create direct exposure and should be removed from the prompt.
The same review should ask what happens when the model does the opposite of each instruction. If ignoring one sentence can release data, approve a payment, grant access, or call a powerful tool, that sentence describes a security control that belongs in deterministic code. The prompt can repeat the intended rule for reliability, but the application must enforce it independently.
Examples embedded in prompts require equal scrutiny. Teams sometimes remove live credentials but leave realistic tokens, customer records, private URLs, or internal identifiers in demonstrations. Examples should use synthetic values and should not reveal enough operational detail to turn prompt disclosure into a practical attack guide.
Brokered capabilities and scoped credentials
The model should receive capabilities rather than reusable secrets. A server-side broker can expose a narrow refund, search, or notification operation, authenticate to the underlying service, validate arguments, enforce policy, and return only the minimum result. The prompt may describe the tool, but it never needs the credential that powers it.
Capabilities should be scoped to the current user, tenant, resource, action, and time window. Short-lived credentials reduce the value of accidental exposure, while server-bound identifiers prevent the model from substituting another account. Read and write operations should remain separate so a tool needed for research cannot silently modify production data.
A capability broker also creates a consistent audit point. It can record the model and prompt version, requesting session, validated arguments, authorization result, approval, downstream identity, and outcome. These records let investigators determine whether the model proposed an unsafe action or an external control failed to stop it.

Prompt lifecycle and change control
System prompts should be versioned, reviewed, and tested as application configuration. A change can alter tool selection, disclosure behavior, refusal boundaries, or the amount of data placed in context. Production releases should identify the prompt version alongside the model, policy, tools, and evaluation set used to approve the change.
Automated scanning can look for credential formats, connection strings, private keys, internal hostnames, and sensitive examples before deployment. Scanning does not prove that a prompt is safe, but it catches common mistakes and creates a repeatable release check. Access to edit prompts should be limited, and emergency changes should receive retrospective review.
Observability should report prompt extraction attempts without treating every request for instructions as a breach. Teams should distinguish curiosity, debugging, injection attempts, and successful disclosure. The more important alert is often what follows: unusual tool use, authorization failures, sensitive queries, or attempts to exploit information learned from the prompt.
Testing the system without trusting the prompt
Security tests should run with the assumption that the prompt is visible and that the model may ignore it. Testers can remove or invert critical instructions, supply direct and indirect prompt injection, request another tenant's data, and ask for unauthorized tool calls. The external controls should reject each attempt for reasons recorded outside the model.
Teams should also test partial disclosure. Even when the model does not reproduce the full prompt, repeated questions may reveal tool names, limits, roles, internal terms, or decision rules. The goal is not to guarantee perfect secrecy. It is to ensure that whatever can be inferred does not provide credentials or bypass deterministic protections.
Implementation checklist
•
Remove credentials, customer data, private endpoints, and enforceable authorization logic from prompts.
•
Treat examples as production content and replace sensitive values with clearly synthetic data.
•
Give the model narrow brokered capabilities instead of reusable credentials or broad service access.
•
Scope capabilities by user, tenant, resource, action, duration, and approval state.
•
Version and scan prompts and test them with their model, tools, policies, and evaluation set.
•
Verify that the application remains safe when the prompt is disclosed or completely ignored.
Evaluating the user and session outside the prompt
External policy can include runtime trust. A valid account may still present elevated risk because it is controlled through an unfamiliar device, suspicious network, automated environment, or behavior that differs from the legitimate user. The prompt cannot reliably assess these signals, and it should not receive raw security data unless the task requires it.
The application can evaluate risk before exposing sensitive data or tools. A low risk session may proceed normally. A session with several concerning signals may be challenged, restricted, or sent for review. The model can receive only the resulting decision or permitted capability, keeping enforcement outside the conversation.
CrossClassify supports this layer through account, device, network, automation, and behavior intelligence. The CrossClassify integration model allows applications to place risk evaluation around signup, login, session activity, and sensitive actions. The system prompt continues to guide the assistant, while external controls decide what the current session may actually do.

Testing for safe prompt disclosure
Security testing should assume that users can learn at least part of the system prompt. Teams can review the prompt for secrets, internal endpoints, sensitive examples, customer information, hidden credentials, and details that make abuse substantially easier. Removing those elements reduces impact even when extraction cannot be fully prevented.
Tests should also verify that the application remains safe when the model ignores the prompt. The model can be instructed to call an unauthorized tool, exceed a limit, access another tenant, reveal protected data, or skip approval. The external systems should reject every attempt for deterministic reasons that can be observed in logs.
Prompt versions still need governance. Changes can alter tool selection, disclosure behavior, or user experience. Teams should review, test, version, and monitor prompts like other application configuration. The important distinction is that prompt governance improves reliability, while real access control and secrets management preserve security.
Conclusion
System prompts are valuable instructions, but they are not vaults. They should shape communication and guide intended behavior without carrying the secrets or sole controls that protect data, money, identity, or infrastructure. A secure system should remain safe even if the prompt becomes visible or the model fails to follow it.
The practical architecture moves credentials into secrets management, authorization into application code or policy engines, tool permissions into narrow interfaces, and sensitive approvals into deterministic or human controlled workflows. It validates every action and records the decision outside the model.
Runtime identity and behavior complete the design by evaluating the session requesting access. CrossClassify can provide that evidence without replacing the application's policy or prompt. This division of responsibility allows the model to remain flexible while the security boundary remains explicit, testable, and enforceable.
See How CrossClassify Secures the New AI Agent Attack Surface
Classify human sessions, trusted agents, and malicious automation before requests turn into risk

Explore CrossClassify today
Detect and prevent fraud in real time
Protect your accounts with AI-driven security
Try CrossClassify for FREE—3 months
Share in
Related articles
Frequently asked questions
Let's Get Started
Create your free
account today
Discover how to secure your app against fraud using CrossClassify
No credit card required



