Engineering posts about Metrics
Curated summaries and key learnings for engineers working with Metrics.
Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA
The Gemini Enterprise Agent Platform has introduced generally available features for agent and model evaluations, allowing developers to measure agent performance against predefined metrics during...
Using observability data to prevent incidents
The article emphasizes the importance of using observability data to transition from reactive incident response to proactive reliability intelligence. It outlines how engineering teams can leverage...
Observability for any agent, anywhere: Production-ready tracing with OpenTelemetry & Unity Catalog on Databricks
The article discusses the challenges of traditional observability tools in managing the massive volumes of trace data generated by AI agents. It presents a solution through Databricks' integration...
Monitoring reliably at scale
The article outlines the challenges of maintaining reliable observability in systems that are heavily dependent on shared infrastructure, such as Kubernetes and service meshes. It highlights the...
Building a fault-tolerant metrics storage system at Airbnb
The article details Airbnb's development of a high-throughput metrics storage system capable of ingesting 50 million samples per second and managing 2.5 petabytes of data. It outlines the challenges...
Building a high-volume metrics pipeline with OpenTelemetry and vmagent
This article outlines a comprehensive approach to migrating a high-volume metrics pipeline from StatsD to OpenTelemetry and Prometheus. It discusses the challenges faced during the migration, such as...