NIST TEVV-Athlon Shifts AI Risk Management From Checklists

NIST TEVV-Athlon Shifts AI Risk Management From Checklists

Desiree Sainthrope is a legal powerhouse who has navigated the shifting sands of global compliance for years, specializing in the intricate interplay between trade agreements and emerging technologies. As a recognized authority on intellectual property and artificial intelligence, she brings a uniquely sharp perspective to the challenges of governance in an era where software can learn and evolve. Today, she shares her insights into the groundbreaking NIST TEVV-Athlon framework, a move away from the rigid, checklist-driven compliance of the past and toward a more nuanced, objective-focused interrogation of AI systems. This conversation explores the critical importance of understanding what an evaluation truly proves, the dangers of letting metrics become targets, and the necessity of embracing “uncomfortable” data to protect organizational integrity. Our discussion highlights the fundamental shift required by leadership to move from demonstrating that a process was simply finished to truly understanding the risks that remain in the shadows of innovation.

Standardized checklists often fail to capture the unique risks of specific AI applications. How should leadership shift their mindset to ensure that governance actually reflects the nuances of the technology they are deploying rather than just satisfying a procedural requirement?

Leadership must recognize that the era of the universal checkbox is effectively over because AI systems and their environments are far too varied for a one-size-fits-all approach to work. Instead of treating governance as a hurdle to jump over, the TEVV-Athlon framework suggests a more organic, four-stage process that starts with the very core of why the AI was built in the first place. This journey begins with the “articulate and organize” phase, where we define our actual objectives, followed by “define and construct,” “apply and measure,” and finally “synthesize and interrogate.” It requires a certain level of intellectual honesty to admit that simply completing an assessment isn’t the same thing as identifying the right risks. We have all seen that natural tendency to standardize everything until the process itself becomes more important than the insights it was meant to provide, essentially making Goodhart’s Law manifest. To avoid this, executives need to ask whether their evaluation was designed to tell them what they actually need to know about a system in its specific, real-world context, rather than just seeking the comfort of a passing grade.

There is a distinct psychological comfort in seeing a “passing” score on a risk assessment, but you’ve noted that this can create a dangerous false sense of security. Can you explain the risks of letting metrics become the ultimate target in AI evaluation?

The danger is that once a specific measurement becomes the target for a team, it immediately ceases to be an effective measure of actual safety or performance. In the high-stakes world of AI, we often see systems that are meticulously optimized to perform brilliantly against a particular benchmark, yet they crumble when faced with the gritty, unpredictable variables of the real world. This “passing score” creates a sense of confidence that can be entirely misplaced if management doesn’t understand exactly what passed and under what narrow conditions. A truly robust evaluation, as NIST recommends, shouldn’t aim to make management feel comfortable; it should aim to produce evidence that highlights exactly where uncertainty and residual risk still live. Sometimes, the most valuable output of a risk assessment is an uncomfortable answer—one that says the AI works beautifully under certain circumstances but that we simply don’t have enough evidence to trust it under others. This transparency is what prevents catastrophic failures down the road, as it allows leadership to make decisions based on reality rather than a sanitized, benchmarked fantasy.

For compliance and risk professionals who may not be data scientists, what are the most critical questions they should be asking when presented with an AI evaluation report, especially when it comes from a third-party vendor?

You certainly don’t need to be a data scientist to challenge the validity of an AI evaluation, but you do need to be a relentless skeptic of the narrative being presented. I suggest focusing on four fundamental questions that can cut through the noise: what were we trying to learn, why were these specific measurements chosen, what assumptions influenced these results, and what was left outside the scope of testing? These questions shift the focus from “did you do the test?” to “what does this evidence actually prove in relation to our business goals?” This is particularly vital when dealing with third-party vendors who might present glowing benchmark results that have very little to do with how your specific organization intends to use the tool. We have to examine the relationship between the data provided and the actual decisions being made, ensuring that we aren’t just inheriting a vendor’s assumptions. By interrogating the “why” and the “what wasn’t tested,” compliance professionals can ensure they are acting as a true bridge between technical evidence and strategic risk management.

Standardization is often the primary goal for auditors and regulators who crave consistency, yet the strength of this new framework lies in its flexibility. How do we balance the need for organizational consistency with the requirement for tailored AI assessments?

This is the great challenge of modern governance: satisfying the hunger for consistency while respecting the complexity of the technology. Auditors want a paper trail they can easily follow, and executives want results they can digest in a five-minute briefing, both of which are reasonable expectations. However, the problem starts when the act of demonstrating that you followed the process becomes more important than understanding what the process told you about your risks. NIST AI 200-2 offers a structured way to determine what you need to know, but its flexibility is actually its greatest asset, not a weakness that needs to be “corrected” by rigid standardization. We have to maintain the integrity of the framework’s stages—the “how”—while allowing the “what” to be customized to the specific AI system in question. If we force every AI through the exact same testing gauntlet, we will inevitably miss the unique vulnerabilities that could lead to significant legal or financial fallout. Consistency should be found in the rigor of the interrogation, not in a static list of questions that might be irrelevant by next Tuesday.

How should organizations approach the “synthesize and interrogate” phase of the TEVV-Athlon framework to ensure that the results are actually useful for executive decision-making?

The final phase of synthesis and interrogation is where the rubber truly meets the road, as it’s the point where raw data is transformed into a narrative that can drive corporate strategy. It is not enough to simply hand over a stack of measurement results; the risk team must interpret what those results mean for the organization’s specific risk appetite and operational goals. This involves looking at the data from multiple angles, using complementary evaluation approaches to see if they all point to the same conclusion or if there are glaring contradictions. We have to be willing to look at the gaps in the data and ask ourselves if we are comfortable moving forward with that level of uncertainty. It is about moving from “the testing is complete” to “here is what we can reasonably conclude about our system’s safety.” This phase requires a deep collaboration between the technical teams who ran the tests and the compliance and legal experts who understand the broader implications of those results. When done correctly, this interrogation provides the clarity needed to make bold, informed decisions about deploying AI in ways that truly benefit the company.

What is your forecast for the future of AI governance over the next two years?

I anticipate that from 2026 to 2028, we will see a major shakeout where the companies still relying on superficial checklists will face significant regulatory and reputational hurdles. The “check-the-box” mentality is quickly becoming a liability rather than a shield, and we will likely see more legal precedents that hold organizations accountable for the quality of their evaluations rather than just the completion of them. We are moving toward a world where “defensible AI” is the gold standard—meaning you can prove not just that your AI works, but that you deeply understand its limitations and have actively managed its residual risks. I also expect to see a rise in independent AI auditing firms that specialize in these adaptable frameworks, providing a level of objective interrogation that internal teams might find difficult to maintain. Ultimately, the winners will be the organizations that treat AI governance as a strategic discipline, using frameworks like TEVV-Athlon to foster a culture of transparency and rigorous inquiry. This shift will turn risk management from a bottleneck into a competitive advantage, allowing firms to innovate with a level of confidence that their peers simply cannot match.

Subscribe to our weekly news digest.

Join now and become a part of our fast-growing community.

Invalid Email Address
Thanks for Subscribing!
We'll be sending you our best soon!
Something went wrong, please try again later