Methods and protocols
First field-pilot cohort
The 0.3 cohort is open to independent professionals, small and medium-sized organizations, nonprofits, and public services in Switzerland and the European Union. It exists to test the playbook in real work. It is not a request for testimonials and it does not assume that AI will improve the workflow.
Two layers in one 0.3 cycle
Each pilot starts from a preregistered extrapolated range or local hypothesis. This first layer supports planning and may come from comparable tasks in other organizations. The second layer is the pilot observation over the complete denominator. The report retains the projection, the result, and their gap so future transfers can improve. Only the observed layer can count as admitted field feedback, but both belong to the 0.3 work.
Use the public pilot intake to propose a non-identifying pilot. The issue contains coordination metadata only. Raw evidence remains in an authorized private system.
Cohort target
Version 0.3 requires at least three admitted reports:
- one non-agentic or copilot workflow, such as retrieval, classification, prediction, conversation, multimodal processing, or assisted generation;
- one bounded A2 business agent that carries a complete workflow while a human retains the defined approval or veto;
- one additional context that makes the three-report cohort cover both Switzerland and the European Union in total, while testing different work, people, sector duties, or operating conditions.
An orchestrated-agency candidate is useful only when a genuine system already exists and can be compared with a simpler design. It is not required for 0.3. Without such a report, agency-level field performance remains unproven.
Three reports are a learning threshold, not statistical validation. They do not create a universal productivity benchmark.
Participation path
- Screen. Open a public intake with an alias, use pattern, territory, bounded workflow, current stage, and decision question.
- Agree the protocol. Name the local evidence owner, independent reviewer, publication authority, and withdrawal route before measurement.
- Freeze the comparison. Record system version, complete workload, baseline, transferred sources, planned range, retained human work, exclusions, cases, thresholds, incidents, and stop conditions.
- Run privately. Keep prompts, logs, screenshots, personal data, client material, secrets, and security-sensitive detail outside GitHub.
- Compare and review. Position the observed result against the range, explain the gap, then check provenance, denominator, failures, anonymization, transfer limits, and consistency with the raw evidence.
- Admit or withhold. Publish only a sanitized report that satisfies every registry rule. Withhold unsafe, unverifiable, or re-identifying material.
Minimum observations
Every pilot records the complete workload and the share that was actually eligible for AI. Common measures are accepted outcomes, human time, major corrections, critical incidents, trace completeness, and implementation effort. The evaluation profile then changes with the use pattern:
| Use pattern | Minimum additional observations |
|---|---|
| Retrieval | answer support, citation correctness, source authorization, freshness |
| Classification or prediction | calibration, abstention, segment errors, drift |
| Conversation | unsupported answers, disclosure, abuse, human handoff |
| Multimodal | quality by modality, file failures, rights, accessibility |
| Business agent | approvals, tool calls, external effects, effect receipts, safe return |
| Orchestrated agency | comparable outcome, coordination failures, loops, cost, specialist isolation |
Public status
- Recruiting region: Switzerland and the European Union
- Target: 3 admitted reports
- Admitted reports: 0
- Public registry:
field-notes/index.json
An intake, an active pilot, or a draft report does not change the admitted count. The count changes only after independent review and publication.
To print or save as PDF: Ctrl+P (⌘P on Mac).
Source and history · GitHub