[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"academy-blogs-en-1-1-all-golang-the-series-ep167-monitoring-ai-performance-all--*":3,"academy-blog-translations-alxa6bnvrti9g4h":95},{"data":4,"page":80,"perPage":80,"totalItems":80,"totalPages":80},[5],{"alt":6,"collectionId":7,"collectionName":8,"content":9,"cover_image":10,"cover_image_path":11,"created":12,"created_by":13,"expand":14,"id":88,"keywords":89,"locale":60,"published_at":90,"scheduled_at":76,"school_blog":84,"short_description":91,"status":82,"title":92,"updated":93,"updated_by":94,"slug":85,"views":87},"Monitoring AI Latency with Prometheus Metrics in Golang","sclblg987654321","school_blog_translations","\u003Cp>Welcome to EP.167! It's been 166 episodes. Our Go-based AI Backend is fully equipped with Multi-LLM support, Redis Caching, Rate Limiting, Load Balancing, and even Circuit Breakers—ready to scale at an enterprise level.\u003C\u002Fp>\u003Cp>But in the production world, there's a classic engineering adage: \u003Cem>\"If you can't measure it, you can't improve it.\"\u003C\u002Fem>\u003C\u002Fp>\u003Cp>I still remember deploying an AI system to production for the first time. During peak traffic, the system inexplicably slowed down. It took hours to hunt down which AI Provider was bottlenecking our app.\u003C\u002Fp>\u003Cp>AI systems present latency challenges vastly different from standard CRUD APIs. A typical API might respond in 50ms, whereas an LLM can take anywhere from 500ms to 10 seconds, depending on prompt size, context length, and output tokens. Without proper \u003Cem>Observability &amp; Monitoring\u003C\u002Fem>, we have no way of knowing:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>Which LLM provider is degrading during peak loads?\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>Is the latency distribution (\u003Cem>Percentiles: p50, p90, p99\u003C\u002Fem>) meeting our SLA requirements?\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>How often are cache hits\u002Fmisses or circuit breaker trips occurring?\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Today, we’re going to attach a \"heart rate monitor\" to our Go application using Prometheus Metrics!\u003C\u002Fp>\u003Ch2>Understanding Prometheus Metric Types\u003C\u002Fh2>\u003Cp>Prometheus offers several metric types tailored for different use cases. For tracking AI performance, we'll focus on three main ones:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>\u003Cstrong>Counter:\u003C\u002Fstrong> A cumulative metric that only goes up (unless the service restarts). Perfect for tracking the total number of requests (\u003Ccode>ai_requests_total\u003C\u002Fcode>) or total token usage.\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>Gauge:\u003C\u002Fstrong> A metric that can arbitrarily go up and down. Ideal for measuring the number of concurrent requests currently being processed (\u003Ccode>ai_concurrent_requests\u003C\u002Fcode>).\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>Histogram:\u003C\u002Fstrong> Samples observations and counts them in configurable \"buckets\" to calculate statistical distributions. This is the absolute best choice for measuring Latency or Processing Duration (\u003Ccode>ai_request_duration_seconds\u003C\u002Fcode>).\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Ch2>Installing Dependencies\u003C\u002Fh2>\u003Cp>We'll be using the official Prometheus Go client:\u003C\u002Fp>\u003Cp>Bash\u003C\u002Fp>\u003Cpre>\u003Ccode>go get github.com\u002Fprometheus\u002Fclient_golang\u002Fprometheus\ngo get github.com\u002Fprometheus\u002Fclient_golang\u002Fprometheus\u002Fpromhttp\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Structuring and Embedding Prometheus Metrics in Go\u003C\u002Fh2>\u003Cp>We will create custom metrics to clock the response times of various AI providers (e.g., OpenAI, Anthropic, Ollama Local) and categorize them by response status (Success \u002F Error).\u003C\u002Fp>\u003Ch3>Part 1: Declaring Prometheus Variables and Mocking the AI Call\u003C\u002Fh3>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>package main\n\nimport (\n\t\"fmt\"\n\t\"log\"\n\t\"math\u002Frand\"\n\t\"net\u002Fhttp\"\n\t\"time\"\n\n\t\"github.com\u002Fprometheus\u002Fclient_golang\u002Fprometheus\"\n\t\"github.com\u002Fprometheus\u002Fclient_golang\u002Fprometheus\u002Fpromauto\"\n\t\"github.com\u002Fprometheus\u002Fclient_golang\u002Fprometheus\u002Fpromhttp\"\n)\n\nvar (\n\t\u002F\u002F 1. Counter: Total requests categorized by provider and status\n\taiRequestsTotal = promauto.NewCounterVec(\n\t\tprometheus.CounterOpts{\n\t\t\tName: \"ai_requests_total\",\n\t\t\tHelp: \"Total number of AI API requests\",\n\t\t},\n\t\t[]string{\"provider\", \"status\"},\n\t)\n\n\t\u002F\u002F 2. Histogram: Processing time (Latency) categorized by provider\n\taiRequestDuration = promauto.NewHistogramVec(\n\t\tprometheus.HistogramOpts{\n\t\t\tName:    \"ai_request_duration_seconds\",\n\t\t\tHelp:    \"AI API processing duration in seconds\",\n\t\t\t\u002F\u002F Define statistical buckets ranging from 0.1s to 10s\n\t\t\tBuckets: []float64{0.1, 0.5, 1.0, 2.0, 5.0, 10.0},\n\t\t},\n\t\t[]string{\"provider\"},\n\t)\n\n\t\u002F\u002F 3. Gauge: Number of concurrent requests currently running\n\taiActiveRequests = promauto.NewGaugeVec(\n\t\tprometheus.GaugeOpts{\n\t\t\tName: \"ai_active_requests\",\n\t\t\tHelp: \"Number of concurrent AI requests currently processing\",\n\t\t},\n\t\t[]string{\"provider\"},\n\t)\n)\n\n\u002F\u002F simulateAICall mocks an AI provider request and measures processing time\nfunc simulateAICall(provider string) (string, error) {\n\t\u002F\u002F Increment active requests in Gauge\n\taiActiveRequests.WithLabelValues(provider).Inc()\n\tdefer aiActiveRequests.WithLabelValues(provider).Dec()\n\n\t\u002F\u002F Start timer\n\tstartTime := time.Now()\n\n\t\u002F\u002F Simulate AI processing time (0.2s to 3.0s)\n\tprocessingTime := time.Duration(200+rand.Intn(2800)) * time.Millisecond\n\ttime.Sleep(processingTime)\n\n\tduration := time.Since(startTime).Seconds()\n\n\t\u002F\u002F Record Latency in Histogram Bucket\n\taiRequestDuration.WithLabelValues(provider).Observe(duration)\n\n\t\u002F\u002F Simulate a 10% chance of error\n\tif rand.Float32() &lt; 0.1 {\n\t\taiRequestsTotal.WithLabelValues(provider, \"error\").Inc()\n\t\treturn \"\", fmt.Errorf(\"AI Provider %s Timeout \u002F Service Unavailable\", provider)\n\t}\n\n\taiRequestsTotal.WithLabelValues(provider, \"success\").Inc()\n\treturn fmt.Sprintf(\"Response from %s (took %.2fs)\", provider, duration), nil\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3>Part 2: Setting up HTTP Endpoints and Prometheus Scraping: main.go\u003C\u002Fh3>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>func main() {\n\t\u002F\u002F 1. Endpoint for Prometheus Server to scrape metrics\n\thttp.Handle(\"\u002Fmetrics\", promhttp.Handler())\n\n\t\u002F\u002F 2. Mock API Endpoint for AI processing\n\thttp.HandleFunc(\"\u002Fapi\u002Fv1\u002Fgenerate\", func(w http.ResponseWriter, r *http.Request) {\n\t\tprovider := r.URL.Query().Get(\"provider\")\n\t\tif provider == \"\" {\n\t\t\tprovider = \"openai\" \u002F\u002F Default Provider\n\t\t}\n\n\t\tresp, err := simulateAICall(provider)\n\t\tif err != nil {\n\t\t\thttp.Error(w, err.Error(), http.StatusInternalServerError)\n\t\t\treturn\n\t\t}\n\n\t\t_, _ = w.Write([]byte(resp))\n\t})\n\n\tlog.Println(\"📊 Prometheus Metrics Endpoint running at http:\u002F\u002Flocalhost:2112\u002Fmetrics\")\n\tlog.Println(\"🚀 AI Service API ready to receive requests at http:\u002F\u002Flocalhost:2112\u002Fapi\u002Fv1\u002Fgenerate\")\n\tlog.Fatal(http.ListenAndServe(\":2112\", nil))\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Building a Grafana Dashboard with PromQL\u003C\u002Fh2>\u003Cp>Once Prometheus starts scraping metrics from our \u003Ccode>\u002Fmetrics\u003C\u002Fcode> endpoint into its time-series database, we can use \u003Cem>PromQL (Prometheus Query Language)\u003C\u002Fem> in Grafana to build sleek, insightful dashboards:\u003C\u002Fp>\u003Ch3>1. Calculating the 95th Percentile Latency (How fast 95% of users get a response)\u003C\u002Fh3>\u003Cpre>\u003Ccode>histogram_quantile(0.95, sum(rate(ai_request_duration_seconds_bucket[5m])) by (le, provider))\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3>2. Calculating the Success Rate (%)\u003C\u002Fh3>\u003Cpre>\u003Ccode>(sum(rate(ai_requests_total{status=\"success\"}[5m])) \u002F sum(rate(ai_requests_total[5m]))) * 100\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>🎯 Daily Mission\u003C\u002Fh2>\u003Cp>Try running the code above and use cURL to send multiple requests, swapping between providers:\u003C\u002Fp>\u003Cp>Bash\u003C\u002Fp>\u003Cpre>\u003Ccode>curl \"http:\u002F\u002Flocalhost:2112\u002Fapi\u002Fv1\u002Fgenerate?provider=openai\"\ncurl \"http:\u002F\u002Flocalhost:2112\u002Fapi\u002Fv1\u002Fgenerate?provider=gemini\"\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>Then, open your browser and navigate to \u003Ca rel=\"noopener noreferrer\" href=\"http:\u002F\u002Flocalhost:2112\u002Fmetrics\">\u003Ccode>http:\u002F\u002Flocalhost:2112\u002Fmetrics\u003C\u002Fcode>\u003C\u002Fa>.\u003C\u002Fp>\u003Cp>\u003Cstrong>Food for Thought:\u003C\u002Fstrong> Try searching for the \u003Ccode>ai_request_duration_seconds_bucket\u003C\u002Fcode> text on the \u003Ccode>\u002Fmetrics\u003C\u002Fcode> page. Notice how the counts in each bucket correlate with the actual processing time? Furthermore, if you wanted to track \u003Cem>\"Token Usage\"\u003C\u002Fem> in the future to calculate daily costs, which metric type would you choose: Counter, Gauge, or Histogram? Try designing the structure yourself!\u003C\u002Fp>\u003Cdiv data-type=\"horizontalRule\">\u003Chr>\u003C\u002Fdiv>\u003Ch2>🙋‍♂️ FAQ\u003C\u002Fh2>\u003Ch3>Why not just use standard logging to track execution time? Why go through the hassle of using Prometheus?\u003C\u002Fh3>\u003Cp>Using standard logs (like \u003Ccode>log.Printf(\"Took %v\", duration)\u003C\u002Fcode>) is easy when starting out. However, when your system handles thousands of requests per second, the log volume becomes massive. Computing real-time averages or P95 from plain text logs is incredibly resource-intensive. Prometheus is purposely built to store Time-Series Data, which is far more space-efficient and drastically faster when querying for dashboard visualizations.\u003C\u002Fp>\u003Ch3>How should I configure Histogram Buckets for typical AI workloads?\u003C\u002Fh3>\u003Cp>AI models inherently take much longer to respond than standard APIs. Setting bucket boundaries at the millisecond level (e.g., \u003Ccode>0.01s\u003C\u002Fcode>, \u003Ccode>0.05s\u003C\u002Fcode>) isn't very useful here. It's best practice to configure buckets that cover the realistic processing span of LLMs, such as \u003Ccode>[0.5, 1.0, 2.0, 5.0, 10.0, 30.0, 60.0]\u003C\u002Fcode> seconds. This ensures you can accommodate even extremely large prompts.\u003C\u002Fp>\u003Ch3>Will frequent scraping by Prometheus slow down my server?\u003C\u002Fh3>\u003Cp>Not at all! The \u003Ccode>\u002Fmetrics\u003C\u002Fcode> endpoint is designed to be extremely lightweight and fast. Configuring Prometheus to scrape every 10 or 15 seconds is a standard best practice and will not negatively impact your server's core performance.\u003C\u002Fp>\u003Cdiv data-type=\"horizontalRule\">\u003Chr>\u003C\u002Fdiv>\u003Ch2>📝 Conclusion\u003C\u002Fh2>\u003Cp>In this article, we learned how to arm our Go backend with robust observability to monitor AI performance using \u003Cstrong>Prometheus Metrics\u003C\u002Fstrong>. We explored the three primary metric types (Counter, Gauge, Histogram), wrote clean Go code to accurately measure latency, and applied advanced PromQL queries for Grafana to calculate enterprise-grade statistics like P95 and Success Rates. Your system is now production-ready and fully equipped to confidently answer the business team when they ask about performance!\u003C\u002Fp>\u003Cp>\u003Cstrong>Coming up next (EP.168):\u003C\u002Fstrong> We have an excellent performance monitoring system in place, but one of the biggest headaches for AI developers is that \"AI does not always return errors as standard HTTP Status Codes.\" Sometimes the API returns a 200 OK, but the actual message inside reads \"Sorry, I cannot fulfill this request,\" or it spits out a malformed JSON that our Go system fails to parse! Next time, we're diving deep into \"Error Handling in AI — Managing and mitigating unexpected AI responses.\" Don't miss it, Gophers!\u003C\u002Fp>\u003Cp>\u003Cstrong>Follow Superdev Academy on all platforms:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>\u003Cstrong>🔵 Facebook: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.facebook.com\u002Fsuperdev.academy.th\">\u003Cstrong>Superdev Academy Thailand\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🎬 YouTube: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.youtube.com\u002F@SuperdevAcademy\">\u003Cstrong>Superdev Academy Channel\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>📸 Instagram: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.instagram.com\u002Fsuperdevacademy\u002F\">\u003Cstrong>@superdevacademy\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🎬 TikTok: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.tiktok.com\u002F@superdevacademy?lang=th-TH\">\u003Cstrong>@superdevacademy\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🌐 Website: \u003C\u002Fstrong>\u003Ca rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fsuperdevacademy.com\">\u003Cstrong>superdevacademy.com\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Cp>\u003C\u002Fp>","54c4q0no7jak_tcj19dqisg.png","https:\u002F\u002Ftwsme-r2.tumwebsme.com\u002Fsclblg987654321\u002F421yuu9p9en7yvi\u002F54c4q0no7jak_tcj19dqisg.png","2026-08-04 04:57:29.323Z","76qprkevbgfdps8",{"keywords":15,"locale":54,"school_blog":64},[16,23,28,33,37,41,46,50],{"collectionId":17,"collectionName":18,"created":19,"created_by":13,"id":20,"name":21,"updated":22,"updated_by":13},"sclkey987654321","school_keywords","2026-03-04 08:20:14.253Z","ah6lvy4x8qe08l5","Golang","2026-06-07 06:45:08.193Z",{"collectionId":17,"collectionName":18,"created":24,"created_by":13,"id":25,"name":26,"updated":27,"updated_by":13},"2026-03-04 08:45:38.382Z","8uz7io97gj0jusq","Prometheus","2026-06-07 06:46:52.110Z",{"collectionId":17,"collectionName":18,"created":29,"created_by":13,"id":30,"name":31,"updated":32,"updated_by":13},"2026-03-04 08:44:11.548Z","ezm3p0vsuowuadd","Latency","2026-06-07 06:46:28.821Z",{"collectionId":17,"collectionName":18,"created":34,"created_by":13,"id":35,"name":36,"updated":34,"updated_by":13},"2026-08-04 04:48:50.963Z","vnovu11bx80yr0h","AI Performance",{"collectionId":17,"collectionName":18,"created":38,"created_by":13,"id":39,"name":40,"updated":38,"updated_by":13},"2026-08-04 04:49:10.806Z","wkf0f2sfoww4cky","Go Backend",{"collectionId":17,"collectionName":18,"created":42,"created_by":13,"id":43,"name":44,"updated":45,"updated_by":13},"2026-03-04 08:44:34.753Z","l1a17htphmxe52a","Observability","2026-06-07 06:46:35.412Z",{"collectionId":17,"collectionName":18,"created":47,"created_by":13,"id":48,"name":49,"updated":47,"updated_by":13},"2026-07-29 04:34:11.928Z","t8naihqgaknjfz4","AI Infrastructure",{"collectionId":17,"collectionName":18,"created":51,"created_by":13,"id":52,"name":53,"updated":51,"updated_by":13},"2026-08-04 04:54:41.814Z","5qwly6pnxl68fgr","Golang Tutorial",{"code":55,"collectionId":56,"collectionName":57,"created":58,"flag":59,"id":60,"is_default":61,"label":62,"updated":63},"en","pbc_1989393366","locales","2026-01-22 11:00:02.726Z","twemoji:flag-united-states","qv9c1llfov2d88z",false,"English","2026-04-10 15:42:46.825Z",{"category":65,"collectionId":66,"collectionName":67,"created":68,"expand":69,"id":84,"slug":85,"updated":86,"views":87},"wqxt7ag2gn7xcmk","pbc_2105096300","school_blogs","2026-08-04 04:50:46.945Z",{"category":70},{"blogIds":71,"collectionId":72,"collectionName":73,"created":74,"created_by":13,"id":65,"image":75,"image_alt":76,"image_path":77,"label":78,"name":79,"priority":80,"publish_at":81,"scheduled_at":76,"status":82,"updated":83,"updated_by":13},[],"sclcatblg987654321","school_category_blogs","2026-03-04 08:33:53.210Z","59ty92ns80w_15oc1implw.png","","https:\u002F\u002Ftwsme-r2.tumwebsme.com\u002Fsclcatblg987654321\u002Fwqxt7ag2gn7xcmk\u002F59ty92ns80w_15oc1implw.png",{"en":79,"th":79},"Golang The Series",1,"2026-03-16 04:39:38.440Z","published","2026-06-07 06:45:03.856Z","alxa6bnvrti9g4h","golang-the-series-ep167-monitoring-ai-performance","2026-08-10 12:00:09.403Z",115,"421yuu9p9en7yvi",[20,25,30,35,39,43,48,52],"2026-08-10 10:11:16.855Z","Learn how to implement observability in your Go backend using Prometheus metrics. We cover measuring latency, percentiles (p95\u002Fp99), and request success rates for AI infrastructure, complete with PromQL examples for Grafana.","Golang The Series EP.167: Monitoring AI Performance & Latency with Prometheus","2026-08-10 10:11:16.856Z","423vhnv3ckczcyn",{"th":85,"en":85}]