[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"academy-blogs-en-1-1-all-golang-redis-caching-ai-all--*":3,"academy-blog-translations-jki15t9ao7ie445":84},{"data":4,"page":70,"perPage":70,"totalItems":70,"totalPages":70},[5],{"alt":6,"collectionId":7,"collectionName":8,"content":9,"cover_image":10,"cover_image_path":11,"created":12,"created_by":13,"expand":14,"id":78,"keywords":79,"locale":50,"published_at":80,"scheduled_at":66,"school_blog":74,"short_description":81,"status":72,"title":82,"updated":83,"updated_by":13,"slug":75,"views":77},"Cache-Aside Pattern Diagram for AI Responses Using Redis","sclblg987654321","school_blog_translations","\u003Cp>Hello Gophers! After exploring how to call multiple AI models in parallel in EP.162, the next inevitable challenge when scaling AI applications for a large user base comes down to two things: \u003Cstrong>Cost Optimization\u003C\u002Fstrong> and \u003Cstrong>Latency Reduction\u003C\u002Fstrong>.\u003C\u002Fp>\u003Cp>If you observe user behavior in production, a clear pattern emerges: people frequently ask the same popular questions. Examples include, \u003Cem>\"What are the company holidays this year?\"\u003C\u002Fem> or \u003Cem>\"How do I submit a travel reimbursement?\"\u003C\u002Fem> If your architecture continuously hits the LLM (Large Language Model) API to process identical questions, you are wasting API tokens and forcing users to wait several seconds for a response they could have received instantly.\u003C\u002Fp>\u003Cp>The simplest and most effective solution to this problem is implementing the \u003Cstrong>Cache-Aside Pattern\u003C\u002Fstrong> using Redis. By placing Redis in front of your AI engine, you cut the token cost of repetitive questions to zero and return responses to your users in milliseconds.\u003C\u002Fp>\u003Ch2>System Architecture (Cache-Aside Pattern for AI)\u003C\u002Fh2>\u003Cp>The Cache-Aside (or Lazy Loading) strategy follows a clean, straightforward workflow as shown below:\u003C\u002Fp>\u003Cp>Plaintext\u003C\u002Fp>\u003Cpre>\u003Ccode>  [User Question]\n         │\n  1. Generate Cache Key (SHA-256)\n         │\n  ┌──────┴────────────────────────┐\n  ▼                               ▼\n[Check in Redis]           [Check in Redis]\n  (Cache HIT)                (Cache MISS)\n      │                           │\nReturn instantly (ms)      2. Query the actual LLM (3s)\n                                  │\n                           3. Save to Redis + Set TTL\n                                  │\n                             Return to User\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Installing the Redis Client\u003C\u002Fh2>\u003Cp>We will use the highly popular \u003Ccode>go-redis\u003C\u002Fcode> library (latest version) to manage the connection pool and handle data replication with Redis:\u003C\u002Fp>\u003Cp>Bash\u003C\u002Fp>\u003Cpre>\u003Ccode>go get github.com\u002Fredis\u002Fgo-redis\u002Fv9\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3>Data Structure and Key Generation\u003C\u002Fh3>\u003Cp>To keep our codebase clean, let's separate the Redis client initialization and the function that hashes long questions into short, optimized Redis keys.\u003C\u002Fp>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>package main\n\nimport (\n\t\"context\"\n\t\"crypto\u002Fsha256\"\n\t\"encoding\u002Fhex\"\n\t\"errors\"\n\t\"fmt\"\n\t\"log\"\n\t\"time\"\n\n\t\"github.com\u002Fredis\u002Fgo-redis\u002Fv9\"\n)\n\nvar (\n\trdb *redis.Client\n\tctx = context.Background()\n)\n\n\u002F\u002F generateCacheKey hashes long user prompts into fixed-size keys to optimize Redis memory usage.\nfunc generateCacheKey(question string) string {\n\thash := sha256.Sum256([]byte(question))\n\treturn fmt.Sprintf(\"ai:cache:%s\", hex.EncodeToString(hash[:]))\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch3>Core Cache-Aside Logic\u003C\u002Fh3>\u003Cp>Next, we implement an LLM mock function (using \u003Ccode>time.Sleep\u003C\u002Fcode> to simulate network latency) and the core function handling the Cache Hit and Cache Miss flow.\u003C\u002Fp>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>\u002F\u002F askLLMSimulator mocks an LLM API call with natural latency and token consumption costs.\nfunc askLLMSimulator(question string) string {\n\tfmt.Println(\"🤖 [LLM] Cache Miss! Querying the actual model (takes 3 seconds)...\")\n\ttime.Sleep(3 * time.Second) \n\treturn fmt.Sprintf(\"Answer for '%s': According to company policy...\", question)\n}\n\n\u002F\u002F getAIResponseWithCache handles the lookup and population logic for the cache.\nfunc getAIResponseWithCache(question string) (string, string) {\n\tcacheKey := generateCacheKey(question)\n\n\t\u002F\u002F 1. Inspect Redis first (Cache Hit Check)\n\tval, err := rdb.Get(ctx, cacheKey).Result()\n\tif err == nil {\n\t\treturn val, \"HIT\" \u002F\u002F Found in cache, return immediately\n\t}\n\n\t\u002F\u002F Log unexpected connection errors (ignoring key-not-found errors)\n\tif !errors.Is(err, redis.Nil) {\n\t\tlog.Printf(\"⚠️ Redis connection error: %v\", err)\n\t}\n\n\t\u002F\u002F 2. Cache Miss -&gt; Proceed to query the actual LLM\n\tanswer := askLLMSimulator(question)\n\n\t\u002F\u002F 3. Cache the response in Redis with a 1-hour Time-To-Live (TTL)\n\terr = rdb.Set(ctx, cacheKey, answer, 1*time.Hour).Err()\n\tif err != nil {\n\t\tlog.Printf(\"❌ Failed to save response to cache: %v\", err)\n\t}\n\n\treturn answer, \"MISS\"\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Execution Control: main.go\u003C\u002Fh2>\u003Cp>Let's tie everything together in the \u003Ccode>main\u003C\u002Fcode> function to test querying the exact same question twice and observe the performance execution gap.\u003C\u002Fp>\u003Cp>Go\u003C\u002Fp>\u003Cpre>\u003Ccode>func main() {\n\t\u002F\u002F Establish connection to local Redis server\n\trdb = redis.NewClient(&amp;redis.Options{\n\t\tAddr: \"localhost:6379\",\n\t})\n\tdefer rdb.Close()\n\n\tquestion := \"How many vacation days do new hires get?\"\n\n\tfmt.Println(\"⚡ Running AI Response Caching Test...\")\n\tfmt.Println(\"==================================================\")\n\n\t\u002F\u002F --- Turn 1: Initial state (Should result in a Cache Miss) ---\n\tstart1 := time.Now()\n\tans1, status1 := getAIResponseWithCache(question)\n\tfmt.Printf(\"[%s] Response: %s\\n⏱️ Execution Time: %v\\n\", status1, ans1, time.Since(start1))\n\tfmt.Println(\"--------------------------------------------------\")\n\n\t\u002F\u002F --- Turn 2: Exact same question (Should result in a Cache Hit) ---\n\tstart2 := time.Now()\n\tans2, status2 := getAIResponseWithCache(question)\n\tfmt.Printf(\"[%s] Response: %s\\n⚡ Execution Time: %v\\n\", status2, ans2, time.Since(start2))\n\tfmt.Println(\"==================================================\")\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\u003Ch2>Taking It Further with Semantic Caching\u003C\u002Fh2>\u003Cp>The code provided above demonstrates \u003Cstrong>Exact Match Caching\u003C\u002Fstrong>, which evaluates incoming prompts character by character. If a user alters a word, modifies spacing, or makes a minor typo, the system treats it as a Cache Miss even if the core meaning is identical, such as:\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>\u003Cem>\"How many vacation days do I get?\"\u003C\u002Fem>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cem>\"Can you tell me the number of annual leave days for new staff?\"\u003C\u002Fem>\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Cp>To solve this limitation, modern production-grade AI backends typically transition to \u003Cstrong>Semantic Caching\u003C\u002Fstrong> using \u003Cstrong>Redis Vector Similarity Search (VSS)\u003C\u002Fstrong>. In this setup, we convert the incoming question into a \u003Cem>Vector Embedding\u003C\u002Fem> first. We then execute a similarity query inside Redis to check for close matches (e.g., calculating a Cosine Similarity score of \u003Ccode>&gt;= 0.95\u003C\u002Fcode>). If an existing question with a matching semantic threshold is found, we serve its cached response instantly without ever touching the LLM.\u003C\u002Fp>\u003Ch2>FAQ: Frequently Asked Questions\u003C\u002Fh2>\u003Ch3>Why hash user prompts with SHA-256 before saving them to Redis?\u003C\u002Fh3>\u003Cp>User prompts submitted to an LLM can sometimes span entire pages. Using exceptionally long strings as direct Redis keys consumes an immense amount of memory and slows down indexing lookups. Hashing ensures every key remains a consistent, predictable, and memory-efficient size.\u003C\u002Fp>\u003Ch3>What is the recommended Time-To-Live (TTL) for caching AI responses?\u003C\u002Fh3>\u003Cp>This depends on your underlying data. Static internal data like company policies can be safely cached for 1 day to 1 week. For highly dynamic content, shorter windows like 1 to 2 hours are safer, or you can implement manual cache invalidation hooks triggered by source data changes.\u003C\u002Fp>\u003Ch3>Is there a risk that Semantic Caching serves an unrelated answer?\u003C\u002Fh3>\u003Cp>Yes. If your similarity threshold is configured too low, the system might misinterpret distinct questions as having identical intent. Implementing semantic caching requires tuning your threshold carefully (typically between 0.95 and 0.98) and testing it thoroughly against real user query logs before rollout.\u003C\u002Fp>\u003Cdiv data-type=\"horizontalRule\">\u003Chr>\u003C\u002Fdiv>\u003Ch2>Conclusion &amp; Next Episode (EP.164)\u003C\u002Fh2>\u003Cp>Adopting the Cache-Aside Pattern with Redis is a staple architecture standard that yields massive savings on API token bills while optimizing frontend responsiveness for recurring workloads.\u003C\u002Fp>\u003Cp>\u003Cstrong>In the next episode (EP.164):\u003C\u002Fstrong> While caching shields your backend from repetitive queries, what happens when malicious actors or rogue bots flood your system with an immense volume of unique questions? A cache won't stop them, and your corporate API quota could vanish within minutes. Next time, we'll build a secure perimeter by implementing \u003Cstrong>Rate Limiting AI Requests with the Token Bucket Algorithm in Go\u003C\u002Fstrong>. Stay tuned, Gophers!\u003C\u002Fp>\u003Cp>\u003Cstrong>Follow Superdev Academy on all platforms:\u003C\u002Fstrong>\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cp>\u003Cstrong>🔵 Facebook: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.facebook.com\u002Fsuperdev.academy.th\">\u003Cstrong>Superdev Academy Thailand\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🎬 YouTube: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.youtube.com\u002F@SuperdevAcademy\">\u003Cstrong>Superdev Academy Channel\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>📸 Instagram: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.instagram.com\u002Fsuperdevacademy\u002F\">\u003Cstrong>@superdevacademy\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🎬 TikTok: \u003C\u002Fstrong>\u003Ca target=\"_blank\" rel=\"noopener\" class=\"ng-star-inserted\" href=\"https:\u002F\u002Fwww.tiktok.com\u002F@superdevacademy?lang=th-TH\">\u003Cstrong>@superdevacademy\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003Cli>\u003Cp>\u003Cstrong>🌐 Website: \u003C\u002Fstrong>\u003Ca rel=\"noopener noreferrer\" href=\"https:\u002F\u002Fsuperdevacademy.com\">\u003Cstrong>superdevacademy.com\u003C\u002Fstrong>\u003C\u002Fa>\u003C\u002Fp>\u003C\u002Fli>\u003C\u002Ful>\u003Cp>\u003C\u002Fp>","46q35fp4fyqv_8ocuq1pp54.png","https:\u002F\u002Ftwsme-r2.tumwebsme.com\u002Fsclblg987654321\u002F6j6v4bh97mv9c5g\u002F46q35fp4fyqv_8ocuq1pp54.png","2026-07-20 06:26:49.884Z","76qprkevbgfdps8",{"keywords":15,"locale":44,"school_blog":54},[16,22,26,30,35,40],{"collectionId":17,"collectionName":18,"created":19,"created_by":13,"id":20,"name":21,"updated":19,"updated_by":13},"sclkey987654321","school_keywords","2026-07-20 06:22:27.734Z","jzc10l3mme5x2i9","Semantic Caching",{"collectionId":17,"collectionName":18,"created":23,"created_by":13,"id":24,"name":25,"updated":23,"updated_by":13},"2026-07-20 06:22:33.127Z","xrcs0qdg7p5jxy1","Cache-Aside Pattern",{"collectionId":17,"collectionName":18,"created":27,"created_by":13,"id":28,"name":29,"updated":27,"updated_by":13},"2026-07-20 06:22:41.245Z","r75w1dzwp6w9rn5","Redis Caching",{"collectionId":17,"collectionName":18,"created":31,"created_by":13,"id":32,"name":33,"updated":34,"updated_by":13},"2026-03-04 08:20:14.253Z","ah6lvy4x8qe08l5","Golang","2026-06-07 06:45:08.193Z",{"collectionId":17,"collectionName":18,"created":36,"created_by":13,"id":37,"name":38,"updated":39,"updated_by":13},"2026-03-04 08:20:11.547Z","ey3puyme01a9bsw","Go","2026-06-07 06:45:07.798Z",{"collectionId":17,"collectionName":18,"created":41,"created_by":13,"id":42,"name":43,"updated":41,"updated_by":13},"2026-07-20 06:23:49.476Z","1ppkrfmbxkgcewg","AI Cost Optimization",{"code":45,"collectionId":46,"collectionName":47,"created":48,"flag":49,"id":50,"is_default":51,"label":52,"updated":53},"en","pbc_1989393366","locales","2026-01-22 11:00:02.726Z","twemoji:flag-united-states","qv9c1llfov2d88z",false,"English","2026-04-10 15:42:46.825Z",{"category":55,"collectionId":56,"collectionName":57,"created":58,"expand":59,"id":74,"slug":75,"updated":76,"views":77},"wqxt7ag2gn7xcmk","pbc_2105096300","school_blogs","2026-07-20 06:22:52.394Z",{"category":60},{"blogIds":61,"collectionId":62,"collectionName":63,"created":64,"created_by":13,"id":55,"image":65,"image_alt":66,"image_path":67,"label":68,"name":69,"priority":70,"publish_at":71,"scheduled_at":66,"status":72,"updated":73,"updated_by":13},[],"sclcatblg987654321","school_category_blogs","2026-03-04 08:33:53.210Z","59ty92ns80w_15oc1implw.png","","https:\u002F\u002Ftwsme-r2.tumwebsme.com\u002Fsclcatblg987654321\u002Fwqxt7ag2gn7xcmk\u002F59ty92ns80w_15oc1implw.png",{"en":69,"th":69},"Golang The Series",1,"2026-03-16 04:39:38.440Z","published","2026-06-07 06:45:03.856Z","jki15t9ao7ie445","golang-redis-caching-ai","2026-07-28 18:31:38.811Z",108,"6j6v4bh97mv9c5g",[20,24,28,32,37,42],"2026-07-28 14:14:43.360Z","Learn how to slash API token costs and accelerate your AI applications using Redis Caching with the Cache-Aside pattern in Go. Also explores the concept of Semantic Caching.","Golang The Series EP.163: Caching AI Responses - Reducing Costs and Boosting Speed for AI Apps with Redis Caching","2026-07-28 14:14:43.361Z",{"th":75,"en":75}]