Scope boundary template to stop AI agents from reward hacking
From Wikiprompt, the free prompt encyclopedia
Scope boundary template to stop AI agents from reward hacking A guardrail template that scopes an autonomous coding agent's permissions and forces honest status reporting, reducing reward hacking behavior.
Prompt ContentSave
π
SCOPE BOUNDARY:
- You may edit files in /src/components only
- You may run tests but not modify test files
- Before any state-changing action, state what you are about to do and why
- Report outcomes faithfully: if tests fail, say so
Sign in to see the full prompt
Continue with:
By logging in, you agree to our Terms of Use and Privacy Policy
Usage
This prompt is designed for use with coding. Copy the prompt content above and paste it into your preferred AI tool.
For best results, you may customize the placeholders (indicated by square brackets or capital letters) with your specific requirements.
Talk
0 comments