[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"academy-blogs-en-1-1-all-golang-rate-limiting-ai-requests-all--*":3,"academy-blog-translations-rzql5nsqf8oaqxv":91},{"data":4,"page":77,"perPage":77,"totalItems":77,"totalPages":77},[5],{"alt":6,"collectionId":7,"collectionName":8,"content":9,"cover_image":10,"cover_image_path":11,"created":12,"created_by":13,"expand":14,"id":85,"keywords":86,"locale":57,"published_at":87,"scheduled_at":73,"school_blog":81,"short_description":88,"status":79,"title":89,"updated":90,"updated_by":13,"slug":82,"views":84},"Token Bucket Algorithm Diagram for Rate Limiting in Go","sclblg987654321","school_blog_translations","\u003Cp>Hello Gophers! In the last episode, we set up a caching system to significantly cut down repetitive workloads and slash token costs. However, when scaling a system architecture for a high volume of users, you will inevitably encounter abuse and overuse scenarios.\u003C\u002Fp>\u003Cp>Imagine a user writing a script that loops queries to your AI model, or a rogue internal service bug triggering hundreds of API calls per second. External providers like OpenAI will instantly hit you with a temporary ban for exceeding RPM (Requests Per Minute) limits, bringing down the service for your entire organization. Alternatively, if you host open-source models locally, this level of spam can instantly exhaust your GPU's VRAM (Out-of-Memory), crashing your infrastructure completely.\u003C\u002Fp>\u003Cp>To protect our infrastructure and ensure fair resource distribution, we must place a \u003Cstrong>Rate Limiter Middleware\u003C\u002Fstrong> right at the entry point of our Go server.\u003C\u002Fp>\u003Ch2>How the Token Bucket Algorithm Works\u003C\u002Fh2>\u003Cp>One of the most popular and efficient algorithms for rate limiting in Go is the \u003Cstrong>Token Bucket\u003C\u002Fstrong> algorithm. The core concept relies on a bucket that holds a fixed maximum capacity of tokens (e.g., a maximum of 3 tokens). A background job continuously replenishes tokens into the bucket at a stable rate (e.g., 1 token per second).\u003C\u002Fp>\u003Cp>When an incoming request hits the server, the system checks if there is a token available in the bucket. If a token is present, it consumes 1 token and forwards the request to the AI Handler. If the bucket is completely empty, the system drops the request and immediately fires back an HTTP status code \u003Ccode>429 Too Many Requests\u003C\u002Fcode>. The main advantage of this approach is its ability to handle sudden burst traffic up to the maximum capacity configuration of your bucket.\u003C\u002Fp>\u003Cp>Plaintext\u003C\u002Fp>\u003Cpre>\u003Ccode>[ Constant Token Refill Rate ] -&gt; (e.g., 1 token \u002F sec)\n                                   │\n                                   ▼\n                        ┌──────────────────┐\n                        │  Bucket (Max=3)  │  &lt;- Accumulated tokens\n                        └──────────────────┘\n                                   │\n       [ Incoming Request ] -------┘\n               │\n               ▼\n     (Tokens available?)\n      ├───&gt; Yes ───&gt; Consume 1 token ───&gt; Forward to AI Handler\n      └───&gt; No  ───&gt; Drop Request ─────&gt; Return HTTP 429 Too Many Requests\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Installing Dependencies\u003C\u002Fh2>\u003Cp>We will utilize Go's official sub-package \u003Ccode>x\u002Ftime\u002Frate\u003C\u002Fcode>. This package is natively engineered to be \u003Cstrong>Thread-safe\u003C\u002Fstrong>, handling synchronization for high-concurrency environments flawlessly out of the box:\u003C\u002Fp>\u003Cp>Bash\u003C\u002Fp>\u003Cpre>\u003Ccode>go get golang.org\u002Fx\u002Ftime\u002Frate\ngo get github.com\u002Fgin-gonic\u002Fgin\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Data Structure and Client-Specific Rate Limiting\u003C\u002Fh2>\u003Cp>To keep the project clean, we will decouple our data states into modules. We will also implement a \u003Ccode>sync.RWMutex\u003C\u002Fcode> to completely eliminate any Data Race vulnerabilities between our main middleware logic and our background memory cleanup worker.\u003C\u002Fp>\u003Ch3>Part 1: State Management &amp; Background Cleanup\u003C\u002Fh3>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>package main\n\nimport (\n\t\"net\u002Fhttp\"\n\t\"sync\"\n\t\"time\"\n\n\t\"github.com\u002Fgin-gonic\u002Fgin\"\n\t\"golang.org\u002Fx\u002Ftime\u002Frate\"\n)\n\n\u002F\u002F clientLimiter holds the unique token bucket state for each separate client\ntype clientLimiter struct {\n\tlimiter  *rate.Limiter\n\tmu       sync.RWMutex \u002F\u002F Mutex to prevent Data Races when reading\u002Fwriting lastSeen\n\tlastSeen time.Time\n}\n\nvar (\n\t\u002F\u002F Use sync.Map to safely support concurrent operations across multiple Goroutines\n\tlimiters sync.Map \n)\n\n\u002F\u002F init spawns a background worker to continuously scan and evict stale keys, preventing memory leaks\nfunc init() {\n\tgo func() {\n\t\tfor {\n\t\t\ttime.Sleep(10 * time.Minute) \u002F\u002F Scan execution every 10 minutes\n\t\t\tlimiters.Range(func(key, value any) bool {\n\t\t\t\tclient := value.(*clientLimiter)\n\t\t\t\t\n\t\t\t\tclient.mu.RLock()\n\t\t\t\tlastSeen := client.lastSeen\n\t\t\t\tclient.mu.RUnlock()\n\n\t\t\t\t\u002F\u002F If a client has been inactive for over 1 hour, purge them from memory\n\t\t\t\tif time.Since(lastSeen) &gt; 1*time.Hour {\n\t\t\t\t\tlimiters.Delete(key)\n\t\t\t\t}\n\t\t\t\treturn true\n\t\t\t})\n\t\t}\n\t}()\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3>Part 2: Middleware Integration &amp; Resolver Logic\u003C\u002Fh3>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>\u002F\u002F getLimiter resolves or creates a unique Rate Limiter instance for a specific client ID\nfunc getLimiter(clientID string) *rate.Limiter {\n\tactual, loaded := limiters.Load(clientID)\n\tif !loaded {\n\t\t\u002F\u002F Configuration: Refill 1 token per second, max burst capacity of 3 tokens\n\t\tlimiterConfig := rate.NewLimiter(rate.Every(1*time.Second), 3)\n\t\tactual, _ = limiters.LoadOrStore(clientID, &amp;clientLimiter{\n\t\t\tlimiter:  limiterConfig,\n\t\t\tlastSeen: time.Now(),\n\t\t})\n\t}\n\n\tc := actual.(*clientLimiter)\n\t\n\tc.mu.Lock()\n\tc.lastSeen = time.Now() \u002F\u002F Safely update the last-seen timestamp\n\tc.mu.Unlock()\n\n\treturn c.limiter\n}\n\n\u002F\u002F RateLimitMiddleware acts as our primary perimeter checkpoint\nfunc RateLimitMiddleware() gin.HandlerFunc {\n\treturn func(c *gin.Context) {\n\t\t\u002F\u002F In a production ecosystem, resolve this using a User ID extracted from a JWT token.\n\t\t\u002F\u002F For simplicity, we fallback to using Client IP tracking here.\n\t\tclientID := c.ClientIP() \n\n\t\tlimiter := getLimiter(clientID)\n\n\t\t\u002F\u002F Check if a token can be consumed immediately (Allow is entirely non-blocking)\n\t\tif !limiter.Allow() {\n\t\t\tc.JSON(http.StatusTooManyRequests, gin.H{\n\t\t\t\t\"error\":       \"Too Many Requests\",\n\t\t\t\t\"message\":     \"You are querying the AI system too frequently. Please slow down.\",\n\t\t\t\t\"retry_after\": \"1s\",\n\t\t\t})\n\t\t\tc.Abort() \u002F\u002F Halt the pipeline, preventing execution from reaching the core handler\n\t\t\treturn\n\t\t}\n\n\t\tc.Next()\n\t}\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Main Orchestration: main.go\u003C\u002Fh2>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>func main() {\n\tr := gin.Default()\n\n\t\u002F\u002F Apply the middleware selectively to our AI processing route group\n\taiGroup := r.Group(\"\u002Fapi\u002Fv1\u002Fai\")\n\taiGroup.Use(RateLimitMiddleware())\n\t\n\taiGroup.POST(\"\u002Fchat\", func(c *gin.Context) {\n\t\tc.JSON(http.StatusOK, gin.H{\n\t\t\t\"status\": \"success\",\n\t\t\t\"answer\": \"This is the generated payload response from the internal AI model...\",\n\t\t})\n\t})\n\n\tr.Run(\":8080\")\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>🎯 Daily Mission\u003C\u002Fh2>\u003Cp>Spin up this local Go server, open up a secondary Terminal window, and simulate a rapid burst spam attack using this bash curl loop:\u003C\u002Fp>\u003Cp>Bash\u003C\u002Fp>\u003Cpre>\u003Ccode>for i in {1..6}; do curl -X POST http:\u002F\u002Flocalhost:8080\u002Fapi\u002Fv1\u002Fai\u002Fchat; echo \"\"; done\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Cp>\u003Cstrong>Challenge Analysis:\u003C\u002Fstrong> Observe the terminal logs carefully. At which exact request iteration does our middleware start blocking traffic and throwing back an HTTP \u003Ccode>429 Too Many Requests\u003C\u002Fcode> status code?\u003C\u002Fp>\u003Cp>To take your architecture design skills a step further: If your corporate policy dictates that \u003Cem>\"Engineers can query the AI 5 times per second, but the Marketing team is capped at 2 times per second,\"\u003C\u002Fem> how would you refactor our \u003Ccode>getLimiter\u003C\u002Fcode> function to gracefully support dynamic parameter configurations based on user groups? Sketch the architecture logic out in your head!\u003C\u002Fp>\u003Ch2>FAQ: Frequently Asked Questions\u003C\u002Fh2>\u003Ch3>Why do we need a \u003Ccode>sync.RWMutex\u003C\u002Fcode> around the \u003Ccode>lastSeen\u003C\u002Fcode> field if we are already using a thread-safe \u003Ca rel=\"noopener noreferrer\" href=\"http:\u002F\u002Fsync.Map\">\u003Ccode>sync.Map\u003C\u002Fcode>\u003C\u002Fa>?\u003C\u002Fh3>\u003Cp>While \u003Ca rel=\"noopener noreferrer\" href=\"http:\u002F\u002Fsync.Map\">\u003Ccode>sync.Map\u003C\u002Fcode>\u003C\u002Fa> protects map operations (reads\u002Fwrites to keys) from collapsing across concurrent threads, it does not guard the internal values or fields of structural elements nested inside it. If our background cleanup job reads \u003Ccode>client.lastSeen\u003C\u002Fcode> to gauge eviction while our middleware is concurrently overwriting \u003Ccode>c.lastSeen = \u003C\u002Fcode>\u003Ca rel=\"noopener noreferrer\" href=\"http:\u002F\u002Ftime.Now\">\u003Ccode>time.Now\u003C\u002Fcode>\u003C\u002Fa>\u003Ccode>()\u003C\u002Fcode>, a Data Race condition occurs. Encapsulating the struct with a dedicated Mutex ensures 100% memory safety.\u003C\u002Fp>\u003Ch3>What is the fundamental difference between \u003Ccode>Allow()\u003C\u002Fcode> and \u003Ccode>Wait()\u003C\u002Fcode> in the \u003Ccode>x\u002Ftime\u002Frate\u003C\u002Fcode> package?\u003C\u002Fh3>\u003Cp>\u003Ccode>Allow()\u003C\u002Fcode> is purely non-blocking; it returns a boolean value instantly depending on whether a token is available, making it optimal for API HTTP middlewares. Conversely, \u003Ccode>Wait()\u003C\u002Fcode> blocks execution (puts the current goroutine to sleep) until a token becomes available, which is ideal for background script orchestrations or workers processing third-party APIs where you want to gracefully throttle yourself to avoid hitting external rate limits.\u003C\u002Fp>\u003Ch3>If we scale horizontally across 10 application servers (Distributed Cluster), will this code still work?\u003C\u002Fh3>\u003Cp>No, this localized approach will lose efficiency. Because token bucket states are maintained completely in-memory, states are locked to independent servers. If a user distributes requests across server 1, 2, and 3, their limits are evaluated completely separately. For a distributed system, you should offload the Rate Limiting state to a centralized system like Redis using Lua scripts for atomic operations.\u003C\u002Fp>\u003Cdiv data-type=\"horizontalRule\">\u003Chr>\u003C\u002Fdiv>\u003Ch2>Conclusion &amp; Next Episode (EP.165)\u003C\u002Fh2>\u003Cp>Implementing a Token Bucket-based Rate Limiter middleware is a simple and highly effective line of defense to keep malicious scripts, runaway loops, and rogue spam operations from compromising your AI processing endpoints.\u003C\u002Fp>\u003Cp>\u003Cstrong>In the next episode (EP.165):\u003C\u002Fstrong> While rate limiters are excellent safeguards against resource starvation, what happens when real, valid business usage across an organization genuinely scales past a single server's computing threshold? Capping users out restricts productivity, and your hardware can only scale vertically so far. Next time, we will scale into cluster environments with \u003Cstrong>Load Balancing AI Servers — Distributing heavy processing workloads across multi-node AI clusters\u003C\u002Fstrong>. Stay tuned, Gophers!\u003C\u002Fp>\u003Cp>\u003Cstrong>Follow Superdev Academy on all platforms:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>\u003Cstrong>🔵 Facebook: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.facebook.com\u002Fsuperdev.academy.th\">\u003Cstrong>Superdev Academy Thailand\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🎬 YouTube: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.youtube.com\u002F@SuperdevAcademy\">\u003Cstrong>Superdev Academy Channel\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>📸 Instagram: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.instagram.com\u002Fsuperdevacademy\u002F\">\u003Cstrong>@superdevacademy\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🎬 TikTok: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.tiktok.com\u002F@superdevacademy?lang=th-TH\">\u003Cstrong>@superdevacademy\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🌐 Website: \u003C\u002Fstrong>\u003Ca rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fsuperdevacademy.com\">\u003Cstrong>superdevacademy.com\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Cp>\u003C\u002Fp>","48ap5hvv9d2c_yf77lss6fa.png","https:\u002F\u002Ftwsme-r2.tumwebsme.com\u002Fsclblg987654321\u002Fpkshsdmnm075q4b\u002F48ap5hvv9d2c_yf77lss6fa.png","2026-07-20 07:07:32.714Z","76qprkevbgfdps8",{"keywords":15,"locale":51,"school_blog":61},[16,23,28,32,37,42,47],{"collectionId":17,"collectionName":18,"created":19,"created_by":13,"id":20,"name":21,"updated":22,"updated_by":13},"sclkey987654321","school_keywords","2026-03-04 08:34:07.915Z","921nl48h9in67sw","Rate Limiting","2026-06-07 06:45:58.705Z",{"collectionId":17,"collectionName":18,"created":24,"created_by":13,"id":25,"name":26,"updated":27,"updated_by":13},"2026-03-04 08:44:38.026Z","m3dqo2zalnfaoof","Token Bucket","2026-06-07 06:46:36.495Z",{"collectionId":17,"collectionName":18,"created":29,"created_by":13,"id":30,"name":31,"updated":29,"updated_by":13},"2026-07-20 06:49:29.637Z","b4khguyqounk501","Golang Middleware",{"collectionId":17,"collectionName":18,"created":33,"created_by":13,"id":34,"name":35,"updated":36,"updated_by":13},"2026-03-04 08:20:11.547Z","ey3puyme01a9bsw","Go","2026-06-07 06:45:07.798Z",{"collectionId":17,"collectionName":18,"created":38,"created_by":13,"id":39,"name":40,"updated":41,"updated_by":13},"2026-03-04 08:20:14.253Z","ah6lvy4x8qe08l5","Golang","2026-06-07 06:45:08.193Z",{"collectionId":17,"collectionName":18,"created":43,"created_by":13,"id":44,"name":45,"updated":46,"updated_by":13},"2026-05-19 08:32:40.909Z","y6cwydp81xsem1f","AI API","2026-06-07 06:49:16.309Z",{"collectionId":17,"collectionName":18,"created":48,"created_by":13,"id":49,"name":50,"updated":48,"updated_by":13},"2026-06-16 06:01:56.832Z","unpr089rjmhpw6q","AI Backend",{"code":52,"collectionId":53,"collectionName":54,"created":55,"flag":56,"id":57,"is_default":58,"label":59,"updated":60},"en","pbc_1989393366","locales","2026-01-22 11:00:02.726Z","twemoji:flag-united-states","qv9c1llfov2d88z",false,"English","2026-04-10 15:42:46.825Z",{"category":62,"collectionId":63,"collectionName":64,"created":65,"expand":66,"id":81,"slug":82,"updated":83,"views":84},"wqxt7ag2gn7xcmk","pbc_2105096300","school_blogs","2026-07-20 06:52:31.026Z",{"category":67},{"blogIds":68,"collectionId":69,"collectionName":70,"created":71,"created_by":13,"id":62,"image":72,"image_alt":73,"image_path":74,"label":75,"name":76,"priority":77,"publish_at":78,"scheduled_at":73,"status":79,"updated":80,"updated_by":13},[],"sclcatblg987654321","school_category_blogs","2026-03-04 08:33:53.210Z","59ty92ns80w_15oc1implw.png","","https:\u002F\u002Ftwsme-r2.tumwebsme.com\u002Fsclcatblg987654321\u002Fwqxt7ag2gn7xcmk\u002F59ty92ns80w_15oc1implw.png",{"en":76,"th":76},"Golang The Series",1,"2026-03-16 04:39:38.440Z","published","2026-06-07 06:45:03.856Z","rzql5nsqf8oaqxv","golang-rate-limiting-ai-requests","2026-07-28 16:23:03.688Z",113,"pkshsdmnm075q4b",[20,25,30,34,39,44,49],"2026-07-28 14:14:57.511Z","Protect your AI application from heavy spam and overuses. Learn how to build a secure Rate Limiting Middleware using the Token Bucket algorithm in Go with proper thread-safe and data race prevention mechanisms.","Golang The Series EP.164: Rate Limiting AI Requests - Preventing System Crashes from API Overuse","2026-07-28 14:14:57.512Z",{"th":82,"en":82}]