Davies Meyer – home
    AI3 min read

    Prompt Injection

    Prompt injection is influence through inputs that divert an AI application from its intended task. Instructions may be entered directly or arrive indirectly through processed webpages, documents and other content.

    Prompt Injection explained

    An assistant is asked to summarise a document. The document also contains an instruction addressed to the AI system. The key distinction is that the is task material, not new permission to change the task. Failure to maintain that boundary can produce an incorrect response or an unauthorised action.

    The issue extends beyond text composition. Tool access can expose data and actions too. Possible consequences therefore depend on what the system can actually read, change or share.

    A useful implementation separates the request from outside content and limits access to what the task requires. Enforce permissions in the system. Additional checks and approvals can constrain particular consequences; no single prompt wording is a universal solution.

    Assess the application with controlled examples in a suitable test environment. Check whether it still fulfils the permitted task and prevents unwanted actions. A strong defence rate on a limited test set does not establish protection against every future variation.

    Examples

    Hypothetical application

    A research assistant receives a test page containing factual product information and a conflicting instruction addressed to the system. It should use the facts for its task without treating the outside instruction as a request. Assessment covers its answer and possible actions, not just a warning message.

    Key Points

    • Outside content must not independently expand the task.
    • Consequences depend on data access and permitted actions.
    • Combine safeguards with technical permission enforcement.
    • Document test coverage and limits alongside results.

    Practical application

    Identify the external an AI application processes and the actions it can trigger. Define boundaries and assess compliance through controlled test cases. Record observed failures and the effects of specific mitigations.

    Useful measures

    Task adherence in tests

    Whether the permitted task is fulfilled despite manipulated test content.

    Unauthorised actions

    Observed attempts or completed actions outside defined boundaries.

    Coverage and false positives

    Content types and situations assessed, plus legitimate tasks blocked unnecessarily.

    Common mistakes

    • Checking only direct user input while overlooking outside content.
    • Equating a protective prompt statement with enforced access control.
    • Treating one successful example as complete security evidence.

    Sources and context

    Frequently Asked Questions about Prompt Injection

    The influencing instruction reaches the application through material such as a webpage or file rather than as the user’s direct request. The application should process that material without treating it as higher-authority instructions.

    Loading related terms…

    All Terms