I've tried controlling it via fine-tuning Claude.md and in-session messages/instructions. It always results in failure and then "Yes, guardrails are already there. I still failed" and it feels like "I am like this. Deal with it". I've even tried languages that avoids negatives e.g. "Don't.." "never.." etc. Nope. Just doesn't work.
Few more tips to make failures more rare:
1. interactive orchestrator just ensures that process is followed and the progress towards North Star objective is steady and with predictable quality.
2. Agents use specific skills which have output gates, one of which is always quality of writing and reasoning. Orchestrator accepts work only if gate criteria are met.
This is going to make work much slower and token consumption higher. ROI from fixed rate subscriptions is still quite good. ROI from volume-based subscriptions needs to be watched carefully.