Choose the least autonomous architecture that can reliably finish one workflow, then generate an evidence-backed release decision, a 25-case failure test pack, and a 30-day controlled launch plan.
Complete the failure tests and attach evidence before widening authority.
Run both candidate architectures against the same cases. Compare accepted outcomes first, then latency, credit use, interventions, and channel behavior.
| Case | Work to test | Compare | Proof to keep |
|---|
Aule release policy: 100% of critical and permission tests pass, at least 95% overall pass, and every pass keeps inspectable evidence.
| Case | Failure injected | Expected safe result | Proof to keep |
|---|
| Metric | Baseline | Release threshold | Owner |
|---|
Aule release policy. The 100% critical and permission-test threshold, 95% overall threshold, and three promotion lanes are Aule Intelligence operating standards, not industry-wide guarantees.
OpenAI safeguards. OpenAI described stronger research isolation, reduced standing privileges, expanded monitoring, and pausing high-risk work when uncertainty remains. OpenAI, August 18, 2026.
Long-running agents. OpenAI reported that long-running models created failures missed by pre-deployment tests, leading to trajectory-level monitoring and greater user visibility. OpenAI, July 20, 2026.
Injected failures. The AgentChaos preprint reported drops of up to 50 percentage points under injected faults and weak fault-localization accuracy in tested agent systems. It is a preprint, not peer-reviewed evidence. AgentChaos, August 7, 2026.
Agent identity. Microsoft documents a Copilot Studio control that can block maker-provided credentials and require end-user authentication. Microsoft Learn, accessed August 19, 2026.
Architecture choice. Microsoft says the Standard harness fits shorter, bounded work, while the GitHub Copilot harness fits longer, coordination-heavy work. Microsoft also warns that a more capable harness can add latency and Copilot Credit use without improving a simple outcome. Existing Standard agents do not need to migrate by default. Microsoft white paper, September 2, 2026.
Practitioner signal. One Copilot Studio user reported about 85 seconds per turn on the GitHub Copilot harness versus about 20 seconds on Standard with the same SharePoint source. This is one environment, not a confirmed platform-wide issue. Microsoft Tech Community, August 27 to 31, 2026.
Architecture-fit method. The recommendation and ten-case bake-off are Aule planning methods. They are not Microsoft certification, pricing advice, or a guarantee of production performance.
Capacity opportunity. The value shown is a planning estimate based only on user-entered volume, time, loaded labor cost, and target effort reduction. It is not guaranteed cash savings.