Moving Beyond Traditional 'LADS' Monitoring for GenAI
These conventional metrics confirm system availability and performance but fail to assess the accuracy or helpfulness of GenAI outputs, requiring a new layer of observability.
Traditional monitoring metrics are insufficient for generative AI, requiring new approaches for cost, safety, and quality.
These conventional metrics confirm system availability and performance but fail to assess the accuracy or helpfulness of GenAI outputs, requiring a new layer of observability.
Determinism, cost, security, and quality represent the four primary domains where generative AI architectures diverge from the predictability of legacy software models.
Unlike traditional applications with predictable outputs, GenAI responses are variable, and their costs fluctuate based on factors like tokens and model choice, making compute expenses dynamic rather than static. New attack vectors, such as prompt injection and PII leakage, bypass conventional error monitoring, while the subjective nature of GenAI outputs requires evaluating relevance and accuracy rather than binary pass/fail outcomes.
This shift necessitates continuous quality evaluation integrated directly into the monitoring stack, as a '200 OK' server response no longer guarantees a helpful or correct answer for the end-user.
The cost structure for Large Language Models (LLMs) is highly dynamic and unpredictable, making it difficult to gain visibility into expenses.
One significant factor is 'token creep,' where increasing context windows from, for example, 4,000 to 32,000 tokens can lead to an eightfold increase in costs if not financially reviewed.
Another issue is 'model drift,' where developers switch to more powerful, expensive models like Opus for quality improvements without fully assessing the financial impact, leading to unexpected budget overruns.
Generative AI applications often incur hidden and redundant costs due to inefficient practices.
Marina Petzel notes that a lack of an effective caching layer can cause the same query to repeatedly hit the API, with research indicating that up to 70% of spend can be redundant because the system pays for expensive model completions multiple times, and switching models can increase costs by as much as 15 times for the same request volume.
Implementing granular tagging is crucial for gaining essential visibility into where money is being spent and controlling costs within generative AI applications.
Feature-level tagging helps attribute costs to specific product areas, which then informs product roadmap decisions and investment priorities. User-level tagging, utilizing User ID or organization names, enables accurate chargebacks to clients or departments and aids in identifying potentially abusive usage patterns.
Additionally, model-level tagging, such as identifying GPT-5.5 or provider names, allows for benchmarking expenses across different models, while endpoint-level tagging tracks costs by region or environment (e.g., production versus staging) for more effective infrastructure planning.
Prompt injection rates, which measure attempts by users to manipulate the model, can be tracked using pattern matching or classifiers with aggressive thresholds. PII detection, identified through methods like regex or named entity recognition, demands a zero-tolerance policy for sensitive data such as social security or credit card numbers in model outputs, regardless of other performance metrics.
Furthermore, content moderation scores are used to identify toxic or biased outputs via specialized toxicity classifiers, while jailbreak attempts involve tracking specific patterns that indicate an override of the system prompt to bypass safety guardrails.
Traditional monitoring metrics like latency, error, traffic, and saturation solely confirm system availability and ignore the actual quality of the AI's output.
A server returning a '200 OK' response can be misleading if the AI model provides a confident yet incorrect or hallucinated answer, effectively lying about its usefulness. Engineers must therefore extend their focus beyond mere system uptime to ensure the reliability and safety of the AI's outputs for end-users.
Hallucination rate tracks claims that are unsupported by grounded data, utilizing both manual reviews and automated checks to identify inaccuracies. Relevance scores measure whether the model's response directly addresses the user's intent, often through techniques like BERT or embedding similarity, while user satisfaction is captured via subjective feedback mechanisms like thumbs up/down, aiming for at least 85% positivity.
Answer completeness uses an LLM-as-a-judge approach to assess if the query was addressed holistically, and RAG (Retrieval-Augmented Generation) quality monitors the efficacy of information retrieval using metrics such as Top K accuracy and normalized discounted cumulative gains.
A modern approach to observability for generative AI involves a layered strategy, where foundational infrastructure signals like latency, error, traffic, and saturation remain necessary.
On top of these traditional metrics, a critical secondary layer addressing cost, safety, and quality is essential for robust GenAI production environments.
Datadog provides agent observability tools specifically designed to track these complex layers, enabling developers to effectively monitor the health of AI models and identify various issues.
Datadog emphasizes that while traditional monitoring of latency, error, traffic, and saturation remains crucial, an additional layer of metrics tailored to generative AI applications is required to ensure their health. This additional layer focuses on cost, safety, and quality, which are increasingly critical for engineers to track.
Datadog offers agent observability products designed to assist in monitoring these complex aspects, enabling users to verify if their applications are running healthy and identify any problems that require fixing. Datadog's solutions help track the comprehensive set of metrics necessary for modern AI systems.
Datadog provides various ways to track application health, ensuring applications are running efficiently and problems are quickly identified. Marina Petzel invites attendees to visit the Datadog booth for more information and offers on their agent observability product.
She concludes by thanking the audience and providing contact information via a QR code.
Answers come from the transcript, with the exact spot cited.
Want the next article from AI Engineer?
When AI Engineer publishes, we'll write it up like the one you just read and email it to you.
AI Engineer published 31 in the last 7 days.
Docker's SBX Sandbox Secures AI AgentsAI Engineer2 hours ago · 11:35 · 12 views · Created 2 hours ago
DatologyAI Generates 12 Trillion Synthetic TokensAI Engineer3 hours ago · 20:31 · 4 views · Created 3 hours ago
AI Agents Accelerate Code, Intensify Debugging BurdenAI Engineer3 hours ago · 18:33 · 2 views · Created 3 hours ago