SecurityBrief UK - Technology news for CISOs & cybersecurity decision-makers
United Kingdom
New Relic launches AI Evaluation for production safety

New Relic launches AI Evaluation for production safety

Wed, 7th Oct 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

New Relic has introduced AI Evaluation, a new feature in its AI Observability platform designed to assess AI application behaviour across the developer-to-production lifecycle.

The launch is part of a broader set of product announcements that also included the Ground Truth command-line interface and Infrastructure 360, both intended to give developers and operations teams more direct access to observability data.

AI Evaluation targets a growing problem as companies move generative AI systems from pilot projects into production. Traditional application monitoring can track latency, uptime and infrastructure performance, but it often cannot show whether an AI model produced a reliable answer, leaked sensitive data or responded to a malicious prompt.

The feature analyses AI performance at the transaction level rather than looking only at individual large language model calls. It links response quality, model behaviour and business impact to distributed traces to help engineers determine whether failures came from prompts, retrieval systems, models or backend infrastructure.

It uses an asynchronous "LLM-as-a-judge" service to scan live telemetry and score issues such as hallucinations, prompt injections and data leaks. Those scores are then attached as attributes to application traces, allowing teams to review technical and AI-related signals in one place.

The approach reflects a wider shift in software operations as enterprises adopt systems that behave less like deterministic code and more like probabilistic services. In practice, that means engineering teams must monitor not only whether an application ran, but also whether an AI-generated response was accurate, safe and worth the cost of producing it.

Tracking AI quality

AI Evaluation also connects quality scoring to compute consumption, allowing teams to compare the output of expensive models with cheaper alternatives. New Relic says this can help organisations judge whether higher model costs are delivering enough improvement in results.

The feature includes pre-built evaluators intended to simplify setup, along with tools for testing prompts before deployment. These include a prompt playground for side-by-side tests, reusable datasets for regression testing and controlled experiments across prompts, models and configurations.

Those functions place the product in a growing market for software designed to govern and monitor generative AI systems. As businesses expand AI use in customer service, software development and internal workflows, vendors are trying to address concerns around reliability, security, compliance and spending.

Broader observability

New Relic's other announcements point to the same effort to embed observability more directly into day-to-day engineering work. The Ground Truth CLI is intended to give developers and AI agents a command-line route into incident investigation and recovery checks, reducing the need to move between several tools. Infrastructure 360, meanwhile, is aimed at platform engineering and operations teams that need a unified view of cloud resources, dependencies and configuration changes.

Together, the three launches suggest New Relic is trying to broaden its role beyond conventional application monitoring. The emphasis is on giving engineering teams a way to connect software behaviour, infrastructure state and AI outputs within a single operational view.

Industry analysts have highlighted the same gap. Existing monitoring systems were built largely for traditional software stacks, where failures could usually be traced through deterministic logs and metrics. Generative AI systems introduce a different class of problem, where an application may be technically available but still produce a poor or unsafe answer.

"Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime," said Brian Emerson, Chief Product Officer, New Relic.

"Unlike point solutions that evaluate isolated LLM calls, we look holistically across the entire application transaction to show exactly what happens whenever AI is involved, giving teams full visibility into technical health, business impact and performance, all within the platform tools that SREs, platform engineers and developers already use," added Emerson.

Stephen Elliot, Group Vice President, I&O, Cloud Operations, and DevOps at IDC, said the shift to production AI systems is changing what enterprises need from monitoring tools.

Enterprise requirements

"As organisations move generative AI applications from early pilots into mission-critical production environments, traditional application performance metrics are no longer sufficient on their own," said Elliot.

"A successful AI implementation requires visibility into both technical health and response quality, including accuracy, safety and model efficiency. Bridging live response evaluation with prompt lifecycle management and full-stack operational telemetry is becoming essential for enterprise engineering and security teams looking to mitigate risk and manage costs effectively," added Elliot.

AI Evaluation will be available in public preview in November.