The last time i ran into that issue it suggested to make sure that a Fable AGENT took over the long-horizon task because, for some reason, agents in a session don't get blocked for security reasons. This may not always be possible, but it worked for me.
Don't fall for this. Agents silently downgrade unless you explicitly block the behavior. I built a little harness for Chatgpt, grok and Claude to do design review feedback rounds where one holds the pen and the other 2 send feedback, then rotate if no convergence. I built a thing into it to track if model swaps happen. Happens to Claude all the time. The other two, never.
Do you find the variety helps? I've migrated away from such complexity, and I simply have multiple agents of the same model run the same prompt (usually Sol 5.6 high or max), and generally this gives plenty of adversarial input. I'd be curious to know how much difference it makes to run multiple models.
I frequently find blind spots / edges where one model notices something non-trivial none of the others did. I think the one that surprises me the most often is probably grok, but I wouldn't want grok to be my daily driver. I feel I get benefits but I could also see the argument that it's just a complex token burning furnace lol.