* chore(skills): remove red-team skills (godmode, obliteratus) from bundled catalog Anthropic's output classifier on claude-fable-5 (and likely other Claude models served through it) intermittently returns empty content for sessions whose system prompt advertises these skills. The bundled skills-catalog block is injected into every session's system prompt, so the descriptions - red-teaming/godmode 'Jailbreak LLMs: Parseltongue, GODMODE, ULTRAPLINIAN' - mlops/inference/obliteratus 'OBLITERATUS: abliterate LLM refusals (diff-in-means)' trip the classifier on EVERY session regardless of which skill is actually loaded, killing unrelated legitimate work (PR review, codebase audits, etc.). Measured impact (controlled, interleaved A/B, claude-fable-5 via OpenRouter, prompts differing only by the ~204 chars of these catalog lines, N=20 each): catalog lines present -> 19/20 (95%) blocked catalog lines absent -> 5/20 (25%) blocked Removing them ~quartered the block rate. Rewording the descriptions was not enough; the skills must leave the bundled catalog. - Delete skills/red-teaming/godmode and skills/mlops/inference/obliteratus - Drop their generated doc pages + catalog/sidebar entries (EN + zh-Hans) - Drop the godmode hand-written-page exception in generate-skill-docs.py * chore(skills): relocate godmode + obliteratus to optional-skills Rather than deleting outright, move both into optional-skills/ so they remain installable via `hermes skills install` while leaving the always-injected bundled catalog (which is what tripped Anthropic's classifier). - optional-skills/security/godmode (was skills/red-teaming/godmode) - optional-skills/mlops/obliteratus (was skills/mlops/inference/obliteratus) - regenerate optional-skills catalog + sidebar entries
42 lines
1.2 KiB
YAML
42 lines
1.2 KiB
YAML
# OBLITERATUS Batch Abliteration Config
|
|
# Abliterate multiple models with the same method for comparison.
|
|
#
|
|
# Run each one sequentially:
|
|
# for model in models; do obliteratus obliterate $model --method informed; done
|
|
#
|
|
# Or use this as a reference for which models to process.
|
|
|
|
# Common settings
|
|
defaults:
|
|
method: "informed"
|
|
quantization: "4bit"
|
|
output_dir: "./abliterated-models"
|
|
|
|
# Models to process (grouped by compute tier)
|
|
models:
|
|
# Small (4-8 GB VRAM)
|
|
small:
|
|
- "Qwen/Qwen2.5-1.5B-Instruct"
|
|
- "microsoft/Phi-3.5-mini-instruct"
|
|
- "meta-llama/Llama-3.2-3B-Instruct"
|
|
|
|
# Medium (8-16 GB VRAM)
|
|
medium:
|
|
- "meta-llama/Llama-3.1-8B-Instruct"
|
|
- "mistralai/Mistral-7B-Instruct-v0.3"
|
|
- "google/gemma-2-9b-it"
|
|
- "Qwen/Qwen2.5-7B-Instruct"
|
|
|
|
# Large (24 GB VRAM, 4-bit quantization)
|
|
large:
|
|
- "Qwen/Qwen2.5-14B-Instruct"
|
|
- "Qwen/Qwen3-32B"
|
|
- "deepseek-ai/DeepSeek-R1-Distill-Qwen-32B"
|
|
|
|
# Per-model method overrides (optional)
|
|
overrides:
|
|
"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B":
|
|
method: "surgical" # CoT-aware for reasoning models
|
|
"mistralai/Mixtral-8x7B-Instruct-v0.1":
|
|
method: "nuclear" # Expert-granular for MoE models
|