prompt-injection-detector ยท diff
git:20260507.50823d8 to git:20260507.44b06c0
1 added, 1 removed. Audit A to A.
---
name: prompt-injection-detector
description: Prompt injection detection and prevention for secure LLM applications
allowed-tools:
- Read
- Write
- Edit
- Bash
- Glob
- Grep
graph:
domains: [domain:software-engineering]
specializations: [specialization:ai-agents-conversational]
- skillAreas: [skill-area:natural-language-processing]
+ skillAreas: [skill-area:hallucination-mitigation-fact-checking, skill-area:safety-redteaming]
roles: [role:ml-engineer, role:backend-engineer]
workflows: [workflow:feature-development, workflow:ml-model-lifecycle]
---
# Prompt Injection Detector Skill
## Capabilities
- Detect prompt injection attempts
- Implement input sanitization
- Configure detection classifiers
- Design defense layers
- Implement canary token detection
- Create injection logging and alerting
## Target Processes
- prompt-injection-defense
- tool-safety-validation
## Implementation Details
### Detection Methods
1. **Pattern Matching**: Known injection patterns
2. **ML Classifiers**: Trained injection detectors
3. **Canary Tokens**: Detect instruction override
4. **LLM-Based**: Use LLM to detect manipulation
5. **Perplexity Analysis**: Unusual input patterns
### Defense Strategies
- Input preprocessing
- Prompt structure design
- Output validation
- Sandboxed execution
- Multi-layer defense
### Configuration Options
- Detection threshold
- Pattern rules
- Classifier model
- Action policies
- Alerting settings
### Best Practices
- Defense in depth
- Regular pattern updates
- Monitor false positives
- Test with red-team inputs
### Dependencies
- rebuff (optional)
- transformers
- Custom classifiers