The race to build better AI has created an equally intense race to measure it.
When a new model launches, benchmark scores often become shorthand for determining which model is the smartest or most capable.
But for financial professionals evaluating AI, there’s a more important question than which model performs best overall: What are they being measured against?
General-purpose AI models are built to perform across an enormous range of subjects and tasks, and their benchmarks reflect that breadth.
But strong performance across broad tests doesn’t necessarily show how well a model can handle the complex, information-heavy tasks financial professionals rely on it to perform.
At Blueflame AI, that distinction shapes how our data science team approaches benchmarking.
Evaluate AI against real financial workflows
Many AI benchmarks essentially consist of sets of questions and expected answers. A model's performance is determined by how many it gets right.
However, the questions being asked matter just as much as the final score. For Blueflame, benchmarking starts with the work real clients are doing on the platform.
Rather than relying solely on broad benchmarks designed to assess general-purpose AI, Blueflame’s data science team develops and adapts evaluations around the challenges AI encounters in financial workflows. That means a model can perform exceptionally well on a general benchmark and still not be the right choice for a particular task.
The goal isn't to identify one universally “best” model. It's to understand which models and systems perform best for the capabilities users actually need.
That's also why Blueflame organizes benchmarks around functionality rather than broad labels.
Instead of assuming a small set of questions can determine whether an AI model is broadly “good at investment banking” or “good at due diligence,” the team evaluates capabilities including metric extraction, multi-hop reasoning, tool calling, refusal, and more.
“We don't look for which AI model is ‘’best,” says Chris Redino, Blueflame AI’s Head of Data Science. “We care which model is best for what users.”
The ability to refuse is a good example of why that distinction matters. AI systems are built to be helpful, but when the information required to answer a question isn't available, producing a convincing response isn't necessarily helpful.
One benchmark Blueflame uses deliberately is a set of questions that cannot be answered from the available information. The system is tested on whether it recognizes that limitation rather than producing an unsupported answer.
Sometimes, the right answer is no answer at all.
Specialized financial workflows can also require specialized benchmarks. Blueflame uses and adapts public benchmarks when they reflect the functionality the team needs to evaluate but creates its own when there is a gap.
Virtual data room (VDR) search is one example. A VDR can contain a large number of documents, many with similar content, while a user may be looking for a highly specific piece of information buried somewhere within them. Blueflame developed a benchmark specifically for that challenge, testing whether the system can answer detailed questions and find the right “needle in the haystack.”
Across these evaluations, Blueflame's data science team has built a suite containing many thousands of questions, with even smaller individual benchmarks containing around 100 questions.
Benchmark the system, not just the model
There is another reason a model's benchmark score can't tell the whole story: the model isn't the entire AI system.
The underlying LLM operates alongside agents, tools, retrieval methods, workflows, algorithms, and different configurations — all of which can affect the result a user receives. That's why Blueflame benchmarks not only foundation models, but the broader systems around them.
In its own testing, the team has seen the same underlying model perform very differently depending on its harness — the surrounding system that determines how the model acts, calls tools, and interacts with other components.
That distinction is particularly relevant as AI systems become increasingly agentic. Asking which foundation model a platform uses is only part of understanding how that platform will perform. How the model retrieves information, uses tools, reasons through a task, and interacts with the rest of the system can be just as consequential.
By benchmarking the full system, Blueflame AI can identify the combination of technologies best suited to each task, not just the highest-scoring model.
Raising the standard for AI in finance
For Blueflame's data science team, benchmarking isn't simply a test performed when a new model comes to market. It's part of the development process itself.
“Benchmarking is the constant,” Chris Redino says.
When the team considers a new methodology, model, or technique, benchmarking provides a way to determine empirically whether that change actually improves performance.
Rather than assuming a newer model or more sophisticated approach is better, the team can test it against the capabilities and workflows that matter.
That process has informed Blueflame's work across areas ranging from RAG and GraphRAG to AI agents. It also creates an ongoing feedback loop: develop, benchmark, understand the results, and use those results to inform what comes next.
Ultimately, that approach reflects a broader priority for AI in finance: accuracy. As Chris puts it, “You could have the lowest latency in the world, but if you get it wrong, no one cares.”
AI models will continue to evolve, and today's highest benchmark score will inevitably be challenged by another model tomorrow. For finance, the more meaningful measure isn't simply where a model ranks. It's whether the complete AI system can accurately and reliably perform the work financial professionals need it to do.
When it comes to AI benchmarking, what you measure against may matter more than who comes out on top.



.png)