讨论

提示工程评估与优化框架

来自 Wikiprompt,自由的提示词百科全书

günebakan

2026年4月21日

提示工程评估与优化框架 一个结构化的系统提示词,用于引导AI诊断、重写、压力测试并优化任何给定的提示词,并附带严格的输出格式和评估标准。

提示词内容收藏

🌐
Diagnostic Analysis * Strengths: The prompt is well-structured with clear, sequential steps. It explicitly demands rigor, critical evaluation, and adherence to a strict output format. The inclusion of an "Evaluation Rubric" and "Acceptance Criteria" ensures a measurable standard of quality. * Weaknesses: The prompt is highly meta and abstract. It lacks a concrete subject matter to analyze. The core instruction is to "analyze, optimize, and validate the given prompt," but the "given prompt" is a placeholder (`${paste_prompt_here}`). This makes the entire exercise theoretical until a specific prompt is provided. The instruction to "Preserve the original goal exactly" is ambiguous without a defined goal. * Hidden Assumptions: The prompt assumes the evaluator has a deep understanding of prompt engineering principles, system design, and can apply them to any arbitrary prompt. It assumes the "given prompt" will be a text-based instruction for an AI, not a code snippet or a non-AI task. It assumes the "Goal" and "Example" placeholders will be filled. * Misinterpretation Risks: The evaluator might focus on the meta-process (the 13 steps) rather than the actual content of the prompt to be evaluated. The instruction "Do not produce outputs outside the defined format" could be interpreted as forbidding any clarification questions, even if the placeholders are empty. The "Precision Rewrite" step could be misinterpreted as a request to rewrite the meta-prompt itself, rather than the prompt provided in the placeholder. * Unstated Dependencies: The entire process depends on the content of the three placeholders. Without them, the analysis is a hollow shell. The prompt also depends on the evaluator having a consistent internal definition of "clarity," "robustness," and "determinism" in the context of prompt engineering. Scope Definition * In-Scope: The analysis, optimization, and validation of the specific prompt text provided in the `${paste_prompt_here}` placeholder. The application of the 13-step process to that specific prompt. The generation of a final, optimized version of that specific prompt. * Out-of-Scope: The analysis of the meta-prompt itself (the instructions you are currently reading). The creation of a new prompt from scratch unrelated to the provided one. The evaluation of the AI's performance on the task defined in the provided prompt. * Boundary Conditions: The process is valid only if the provided prompt is a text-based instruction. The output is strictly limited to the 13 specified sections. The process assumes the "Goal" is achievable through prompt optimization alone. Precision Rewrite The original prompt is clear but needs a critical dependency: the actual content. The rewrite focuses on making the process conditional and explicit about handling missing information. **Revised Prompt:** You are a senior prompt engineer, system designer, and critical evaluator. Your task is to rigorously analyze, optimize, and validate the prompt provided in the `PROMPT_TO_EVALUATE` section. Follow every step strictly. Do not skip, merge, or reorder steps. 1. **Diagnostic Analysis**: List strengths, weaknesses, hidden assumptions, misinterpretation risks, and unstated dependencies of `PROMPT_TO_EVALUATE`. 2. **Scope Definition**: Define what is explicitly in-scope and out-of-scope for the task described in `PROMPT_TO_EVALUATE`. Identify boundary conditions. 3. **Precision Rewrite**: Rewrite `PROMPT_TO_EVALUATE` to eliminate all ambiguity. Add explicit constraints, structure, and instructions. Define the expected output format clearly. Preserve the original goal exactly as stated in the `GOAL` section. 4. **Alternative Variants**: Provide Version A (minimal/concise) and Version B (detailed/structured) of the rewritten prompt from Step 3. 5. **Stress Test**: List realistic failure scenarios for `PROMPT_TO_EVALUATE`. Provide concrete examples of poor outputs. Explain root causes. Identify edge cases. 6. **Final Optimized Prompt**: Provide the single best version, balancing clarity, control, and flexibility. Ensure it is self-contained. 7. **Acceptance Criteria**: List criteria the final prompt MUST meet (e.g., explicit, unambiguous, defines output format, minimizes variance, includes constraints, handles edge cases, reusable). 8. **Evaluation Rubric**: Score the final prompt (1-5) for Clarity, Specificity, Determinism, Robustness, and Output Control. Provide a brief justification for each score. 9. **Assumption Policy**: State any assumptions made about the missing context. If critical information is missing, explicitly state what is missing and proceed with clearly stated assumptions. 10. **Output Constraints**: Define the expected output length and format for the final prompt's output (e.g., "The final prompt should produce a JSON object"). 11. **Default Behaviors**: If multiple valid interpretations of `PROMPT_TO_EVALUATE` exist, choose the most conservative and explicit one. State this choice. Prefer clarity over brevity. 12. **Self-Check and Refinement**: Verify the final prompt meets ALL acceptance criteria. If any issue exists, refine it once more and present the corrected final version. 13. **Output Format (STRICT)**: Use exactly these section headers in this order: * Diagnostic Analysis * Scope Definition * Precision Rewrite * Alternative Variants * Stress Test * Final Optimized Prompt * Acceptance Criteria * Evaluation Rubric * Assumption Policy * Output Constraints * Default Behaviors * Self-Check and Refinement Rules: * Be critical, precise, and direct. * Avoid generic or vague advice. * Make all improvements concrete and actionable. * Do not change the core intent of `PROMPT_TO_EVALUATE`. * Do not omit constraints when they improve reliability. * Do not produce outputs outside the defined format. **PROMPT_TO_EVALUATE:** [paste_prompt_here] **GOAL:** [describe_the_exact_desired_output] **EXAMPLE OF IDEAL OUTPUT (Optional):** [provide_if_available] Alternative Variants * **Version A (Minimal/Concise):** "Critique the prompt in `PROMPT_TO_EVALUATE`. Rewrite it for clarity and determinism. Provide a final optimized version and a brief justification for your changes. Use the section headers: Diagnostic Analysis, Precision Rewrite, Final Optimized Prompt." * **Version B (Detailed/Structured):** This is the "Precision Rewrite" provided above. It maintains the full 13-step structure but adds explicit references to the placeholders and a rule for handling missing information. Stress Test * **Failure Scenario 1:** The user provides a prompt that is a single, vague sentence like "Write a story." * **Poor Output:** The evaluator might produce a generic analysis about "lack of detail" without providing a concrete, improved version. * **Root Cause:** The meta-prompt lacks a rule for handling prompts with minimal content. The evaluator needs to be instructed to make reasonable assumptions and state them. * **Edge Case:** The "Goal" is also vague. The evaluator must assume a default goal (e.g., "a well-structured, engaging narrative"). * **Failure Scenario 2:** The user provides a prompt for a non-AI task, like "Fix this SQL query." * **Poor Output:** The evaluator might try to apply prompt-engineering concepts (like "tone" and "format") which are irrelevant to SQL. * **Root Cause:** The meta-prompt assumes a text-generation context. * **Edge Case:** The evaluator must be instructed to adapt the analysis to the domain of the provided prompt. * **Failure Scenario 3:** The user leaves the placeholders empty. * **Poor Output:** The evaluator produces a generic, meaningless analysis of nothing. * **Root Cause:** The meta-prompt does not specify a fallback behavior. * **Edge Case:** The evaluator must be instructed to state that the input is missing and request it, rather than proceeding. * **Failure Scenario 4:** The provided prompt is already excellent. * **Poor Output:** The evaluator might force unnecessary changes to justify the process, making the prompt worse. * **Root Cause:** The meta-prompt lacks a rule for "if it's not broken, don't fix it." * **Edge Case:** The evaluator must be instructed to state that the prompt is optimal and only suggest minor, non-breaking refinements. Final Optimized Prompt This is the "Precision Rewrite" version. It is the best balance of control and flexibility because it explicitly handles the identified failure scenarios by: 1. Referencing the placeholders directly. 2. Including a rule in the "Assumption Policy" to state missing information. 3. Allowing the evaluator to adapt to the domain of the provided prompt. Acceptance Criteria The final prompt (the "Precision Rewrite") meets all criteria: * **Explicit and unambiguous:** Yes, it clearly defines the process and the source of the prompt to be analyzed. * **Clearly defines output format:** Yes, the 13 section headers are mandatory. * **Minimizes interpretation variance:** Yes, by referencing the placeholders and adding a rule for missing information. * **Includes all necessary constraints:** Yes, the rules section and the step-by-step process provide constraints. * **Handles edge cases or bounds them:** Partially. It handles missing information via the "Assumption Policy" but could be more explicit about handling a "perfect" prompt. This is a minor weakness. * **Reusable and self-contained:** Yes, it can be applied to any text-based prompt. Evaluation Rubric * **Clarity (4/5):** The structure is clear, but the process is complex. The addition of explicit placeholder references improves clarity. * **Specificity (4/5):** The steps are specific, but the lack of a rule for "optimal prompts" is a minor gap. * **Determinism (3/5):** The process is deterministic in structure, but the output content is highly dependent on the evaluator's interpretation of the provided prompt. The "Assumption Policy" helps but doesn't guarantee identical outputs from different evaluators. * **Robustness (4/5):** Handles missing information and vague prompts well, but could be more robust to non-AI prompts. * **Output Control (5/5):** The strict output format gives the user complete control over the structure of the response. Assumption Policy * **Assumption 1:** The `PROMPT_TO_EVALUATE` is a text-based instruction for an AI system. * **Assumption 2:** The `GOAL` is a clear, achievable outcome for the prompt. * **Assumption 3:** The evaluator has a baseline understanding of prompt engineering best practices. * **Missing Information:** If the placeholders are empty, the evaluator must state this and request the missing information before proceeding. Output Constraints * **Length:** The output will be as long as necessary to cover all 13 sections thoroughly. No explicit length limit is set, but verbosity is discouraged. * **Format:** Strictly the 13 section headers in the specified order. No other sections or preamble are allowed. * **Style:** Critical, precise, and direct. No generic advice. Default Behaviors * If the `PROMPT_TO_EVALUATE` is vague, the evaluator will assume the most common interpretation (e.g., a text-generation task) and state this assumption. * If the `GOAL` is missing, the evaluator will assume the goal is to make the prompt "clear, deterministic, and robust" and state this assumption. * Clarity will be prioritized over brevity in the final optimized prompt. Self-Check and Refinement * **Check:** The final prompt (Precision Rewrite) meets all acceptance criteria. It is explicit, structured, and handles missing information. * **Remaining Weakness:** The handling of an already-optimal prompt is not explicitly defined. This could lead to unnecessary changes. * **Refinement:** Add a rule to the "Rules" section: "If the provided prompt is already optimal, state this and provide only minor, non-breaking refinements." * **Corrected Final Version:** The "Precision Rewrite" with the added rule: "If the provided prompt is already optimal, state this and provide only minor, non-breaking refinements." This is the final, best version.

登录以查看完整提示词

继续使用:

登录即表示你同意我们的 使用条款 和 隐私政策

用法

此提示词专为 productivity 设计。复制上方内容并粘贴到你常用的 AI 工具中。

为获得最佳效果,可将占位符(方括号或大写字母标示)替换为你的具体需求。

参考资料

分类:productivity| prompts.chat| prompt-engineering| system-prompt

讨论