You don’t need to resolve the research debate around emergence to manage the risks.
You need clear ownership, measurable standards, and a process for evaluating changes.
Six practical actions can establish that foundation.
1. Map where AI behavior can change
Identify AI-enabled applications, vendor-managed models, fine-tuned models, and agent workflows currently in use.
Document where vendors can introduce model updates, where employees can change prompts or configurations, and where systems can gain access to additional data or tools.
Suggested owner: CIO, CTO, or designated technology leader.
2. Require testing when systems change
Ask AI vendors how they notify customers of model updates, what testing they perform, and what options exist to manage or delay changes.
Where appropriate, establish change-notification and testing expectations in vendor agreements.
For internally managed systems, define which changes trigger a new evaluation.
3. Test beyond the intended use case
When your organization fine-tunes or customizes a model, evaluate it beyond the task you’re trying to improve.
Test for unexpected outputs, inappropriate behavior, security concerns, and failures in adjacent workflows.
The research on emergent misalignment demonstrates why narrowly focused testing may not be sufficient.
4. Evaluate AI agent workflows as complete systems
Before multiple agents operate together in production, test the end-to-end workflow.
Look at how information moves between agents, which actions they can take, what happens when one produces an incorrect result, and whether errors can compound.
Individual component testing remains important, but it should not replace system-level evaluation.
5. Define when human review is required
Identify which decisions AI can support, which actions it can perform independently, and which require human approval.
Establish escalation criteria for errors, unusual behavior, sensitive information, and high-impact decisions.
Make sure the people responsible for oversight have the authority to intervene.
6. Measure AI against business outcomes
Evaluate AI using the standards your business process actually requires.
An invoice-matching system should be measured on accuracy, exception handling, and processing efficiency. A document-summary tool should be evaluated on factual reliability and the consequences of omissions.
Vendor benchmark performance may be useful context, but it is not your business outcome.
How to frame the decision with leadership
We don’t need to eliminate every uncertainty before using AI. We need clear ownership, measurable performance standards, and controls that allow us to detect and respond when behavior changes.
Underneath all six actions sits one essential decision:
Name the person accountable for each AI system’s performance and the process for responding when it falls outside acceptable limits.
That creates clarity for the technology team and confidence for leadership.