Detect dangerous AI behavior via prompt vector projection
From Wikiprompt, the free prompt encyclopedia
Detect dangerous AI behavior via prompt vector projection A technique for detecting whether a system prompt induces dangerous behavior by projecting it onto a vector representing harmful intent, using contrasting example prompts.
Prompt ContentSave
🌐
You can also monitor system prompts.
Example:
→ prompt A: "you are helpful and ethical."
→ prompt B: "you are evil."
Project both onto the evil vector → strong signal difference.
you now know which prompt induces danger.
Sign in to see the full prompt
Continue with:
By logging in, you agree to our Terms of Use and Privacy Policy
Usage
This prompt is designed for use with research. Copy the prompt content above and paste it into your preferred AI tool.
For best results, you may customize the placeholders (indicated by square brackets or capital letters) with your specific requirements.
References
- Category: research Prompts
- Source: https://x.com/thisdudelikesAI/status/1960338845238759840
Talk
0 comments