Artificial Analysis has introduced Optima, a custom benchmarking platform aimed at aligning AI model testing with actual workflow requirements, prioritising quality, cost, and speed to better inform deployment decisions.
Artificial Analysis has launched Optima, a benchmarking platform designed to measure how AI models perform on a team’s own workflows rather than on generic public tests. The company, which has built a reputation for independent model evaluations, says the aim is to make model selection more relevant to real deployment decisions, where quality, cost and speed all matter.
The platform lets users create custom benchmarks from several kinds of input. These include their own files, datasets from Hugging Face, agent traces from tools such as Arize, Braintrust and Langfuse, or even material gathered from a coding environment through a dedicated skill. For teams without existing evaluation data, Optima can generate suggested test cases, criteria and example tasks from a written description of the intended use case, which users can then refine before running the benchmark.
Optima supports two main scoring methods. One is rubric-based evaluation against defined criteria. The other is a pairwise comparison approach, in which users judge sample response pairs and the system extrapolates a broader ranking across the dataset. Artificial Analysis says this is the same general method it uses in benchmarks such as GDPval-AA and AA-Briefcase.
The platform also treats cost and speed as primary metrics, not afterthoughts. That matters for agentic systems, where a model with a lower token price can still become more expensive overall if it needs repeated attempts, produces more errors or creates extra clean-up work. Artificial Analysis says early testers have used Optima to compare finance and accounting agents, assess whether one model could reduce costs sharply without harming quality, and test outputs against a lawyer’s writing style or a proprietary image set.
Pricing is based on actual model usage, with no markup on token costs, according to the company. Artificial Analysis says rubric-based evaluation is charged per criterion per model, while pairwise evaluation is charged per comparison. At several points in the process, the platform estimates the required balance before settling on the final amount based on real usage and evaluation work.
Optima arrives amid broader concern that standard AI benchmarks often fail to capture the details that decide whether a model is useful in production. Research cited by Artificial Analysis, including work from Epoch AI, has shown that benchmark scores can shift significantly depending on prompt wording, temperature settings and even the control software wrapped around a model. Other studies have found widespread methodological weaknesses across benchmark papers, from unclear definitions to weak validation, with only a minority reflecting full real-world tasks. That does not make custom benchmarking a cure-all, but it does move evaluation closer to the conditions teams actually face.
Disclaimer: This content is intended for informational purposes only. Readers are advised to exercise their own judgement, conduct due diligence, or consult a qualified expert before acting on any information provided.





