Contact
Line : comsiam
Contact
Line : comsiam

Context Caching คือความสามารถของ Gemini API ที่ช่วยลดต้นทุนเมื่อ Application ต้องส่งข้อมูลเดิมขนาดใหญ่ให้โมเดลซ้ำหลายครั้ง เช่น System Instruction ยาว ๆ, PDF, วิดีโอ, Code Repository หรือ Conversation History
ตัวอย่าง หากมีเอกสารขนาด 100,000 Tokens และต้องถามเอกสารเดิม 100 ครั้ง หากไม่มี Caching ระบบอาจต้องประมวลผล Context เดิมซ้ำจำนวนมาก
แนวคิดคือ
เอกสารใหญ่
100,000 tokens
↓
คำถามที่ 1
ส่ง Context ทั้งหมด
คำถามที่ 2
ส่ง Context ทั้งหมดอีก
คำถามที่ 3
ส่ง Context ทั้งหมดอีก
...
Context Caching ช่วยให้ Gemini สามารถใช้ Context ที่เคยประมวลผลไว้ซ้ำในราคาที่ต่ำกว่าการคิด Input Token ปกติเมื่อเกิด Cache Hit
ปัจจุบัน Gemini API มี Caching 2 รูปแบบสำคัญคือ
Implicit Caching
และ
Explicit Caching
แต่มีข้อแตกต่างสำคัญมาก:
Interactions API รุ่นปัจจุบันรองรับเฉพาะ Implicit Caching เท่านั้น
หากต้องการสร้าง Cache Object เอง กำหนด TTL และนำ Cache ID กลับมาใช้แบบ Explicit Caching ต้องใช้ generateContent API
ดังนั้นสำหรับ Project ใหม่ที่ใช้ Interactions API ควรเริ่มจาก Implicit Caching ก่อน เพราะเปิดให้อัตโนมัติและไม่ต้องเขียนระบบ Cache เพิ่ม
ลองสมมติว่าเรามีคู่มือ
company-manual.pdf
เมื่อแปลงเป็น Context แล้วมี
100,000 tokens
User ถาม
คำถามที่ 1:
นโยบายคืนเงินคืออะไร
จากนั้นถาม
คำถามที่ 2:
รับประกันสินค้ากี่ปี
และ
คำถามที่ 3:
มีเงื่อนไขอะไรบ้าง
ข้อมูลต้นทางยังเป็น PDF เดิม
หากระบบประมวลผล Context เดิมซ้ำทุกครั้ง ต้นทุน Input จะเพิ่มตามจำนวน Request
Caching ช่วยให้ Context ที่ซ้ำมีโอกาสถูกคิดในอัตรา Cached Input ที่ต่ำกว่า
แนวคิดพื้นฐานคือ
Input ปกติ
=
ราคา Input Token เต็ม
ขณะที่ Cache Hit
Cached Input
=
ราคา Context Caching
ซึ่งในหลาย Model ต่ำกว่า Input ปกติประมาณมาก
ตัวอย่าง Gemini 3.7 Flash Standard Paid Tier ณ วันที่ 2 กันยายน 2026
Input ปกติ
$0.75 / 1M tokens
ขณะที่ Context Caching
$0.075 / 1M cached tokens
จนถึงวันที่ 31 ธันวาคม 2026
เท่ากับ Cached Input Price ประมาณหนึ่งในสิบของ Input Price ในช่วงราคานี้
แต่ Explicit Cache ยังมี Storage Cost ตามระยะเวลาที่เก็บ Cache ด้วย
Standard Pricing ปัจจุบันคือ
| รายการ | ถึง 31 ธ.ค. 2026 | ตั้งแต่ 1 ม.ค. 2027 |
|---|---|---|
| Input | $0.75 / 1M tokens | $1.50 / 1M tokens |
| Context Caching | $0.075 / 1M tokens | $0.15 / 1M tokens |
| Cache Storage | $0.50 / 1M tokens/ชั่วโมง | $1.00 / 1M tokens/ชั่วโมง |
ราคาสามารถเปลี่ยนได้ จึงควรตรวจ Pricing ก่อนวาง Budget Production
สมมติ Context
100,000 tokens
ใช้กับ
1,000 requests
Context Input รวมประมาณ
100,000 × 1,000
=
100,000,000 tokens
หรือ
100 million tokens
ถ้าคิดด้วย Input Rate
$0.75 / 1M tokens
เฉพาะ Context ซ้ำส่วนนี้จะประมาณ
100 × $0.75
=
$75
ยังไม่รวม
สมมติ Context เดิมถูก Cache และ Request ต่อมามี Cached Tokens รวม
100 million tokens
ที่ราคา
$0.075 / 1M
ค่า Cached Input จะประมาณ
100 × $0.075
=
$7.50
แทน $75 สำหรับ Input ส่วนเดียวกัน
แต่ถ้าเป็น Explicit Cache ต้องเพิ่ม Storage Cost เข้าไปด้วย
ดังนั้นสูตรที่ถูกคือ
Total Cache Cost
=
Cached Token Usage
+
Storage
+
New Input
+
Output
+
Tools
ไม่ใช่ดู Cached Token Price อย่างเดียว
Implicit Caching คือระบบ Cache ที่ Gemini จัดการให้อัตโนมัติ
Developer ไม่ต้องสร้าง
Cache Object
ไม่ต้องกำหนด
TTL
และไม่ต้องส่ง
cached_content ID
Google เปิด Implicit Caching ให้อัตโนมัติกับ Gemini 2.5 และใหม่กว่า
เมื่อ Request ตรงกับ Cache ที่ระบบสามารถใช้ได้ Google จะส่ง Cost Saving ให้โดยอัตโนมัติ
ถ้าใช้
client.interactions.create(...)
ไม่ต้องเขียน
enable_cache=True
หรือ Parameter พิเศษเพื่อเปิด Cache
Implicit Caching เปิดอยู่แล้วตาม Model ที่รองรับ
ตัวอย่าง
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.7-flash",
input="อธิบายข้อมูลนี้",
)
print(interaction.output_text)
ระบบจะจัดการ Implicit Cache เอง
คำว่า
Implicit
หมายถึง Google เป็นผู้ตัดสินใจ Cache
Developer ไม่สามารถรับประกันได้ว่า Request ทุกครั้งจะเกิด
Cache Hit
ดังนั้นไม่ควรสร้าง Financial Model ว่า
ทุก Request
=
Cached Price แน่นอน
แต่ควรวัด Cache Hit จาก Usage จริง
Minimum Token สำหรับ Implicit Caching แตกต่างตาม Model
ปัจจุบันตัวอย่างสำคัญคือ
| Model | Minimum Input Tokens |
|---|---|
| Gemini 3.7 Flash | 4,096 |
| Gemini 3.6 Flash | 4,096 |
| Gemini 3.5 Flash | 4,096 |
| Gemini 3.1 Pro Preview | 4,096 |
| Gemini 2.5 Flash | 2,048 |
| Gemini 2.5 Pro | 2,048 |
ถ้า Prompt มีเพียง
100 tokens
อย่าคาดหวัง Context Caching ตาม Minimum Requirement ของ Model
Google แนะนำหลักสำคัญ 2 ข้อ
เช่น
System Instruction
+
Large Common Context
+
User-specific Question
แทน
User-specific Content
+
Random Data
+
Common Context
Cache Matching อาศัยความคล้ายของ Prefix
ดังนั้น Structure ของ Prompt มีผลมาก
สมมติทุก Request มีข้อมูลใหญ่เหมือนกัน
คู่มือบริษัท 50,000 tokens
และมีคำถามแตกต่างกันท้าย Prompt
รูปแบบที่เหมาะคือ
[คู่มือบริษัทเหมือนเดิม]
+
คำถาม A
[คู่มือบริษัทเหมือนเดิม]
+
คำถาม B
[คู่มือบริษัทเหมือนเดิม]
+
คำถาม C
Common Prefix มีโอกาสถูก Cache ได้ดีกว่า
ตัวอย่าง Request 1
Timestamp: 10:01:01
คู่มือบริษัท...
Request 2
Timestamp: 10:01:03
คู่มือบริษัท...
แม้คู่มือจะเหมือนกัน แต่ Prefix เปลี่ยนตั้งแต่ส่วนต้น
อาจลดโอกาส Cache Hit
ดังนั้น Dynamic Data ควรอยู่ส่วนหลังเมื่อไม่จำเป็นต้องอยู่ต้น Prompt
ใช้
System Instruction เดิม
↓
Document เดิม
↓
Static Context
↓
Dynamic User Data
↓
User Question
ทำให้ส่วนที่ซ้ำอยู่ด้านหน้า
แนวคิดนี้เรียกว่า
Prefix Stability
ซึ่งสำคัญต่อ Implicit Cache
Interactions API ปัจจุบันสามารถดูจำนวน Token ที่ Cache Hit ได้จาก Usage
Field คือ
usage.total_cached_tokens
ใน Python และ JavaScript
ดังนั้นอย่าเดาว่า Cache ทำงานหรือไม่
ควรดู Usage จริง
แนวคิดคือ
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.7-flash",
input=large_prompt,
)
print(interaction.usage)
จากนั้นดู
total_cached_tokens
หากมีค่ามากกว่า 0 แสดงว่ามี Input Tokens บางส่วนเกิด Cache Hit
ตัวอย่างแนวคิด
const interaction =
await ai.interactions.create({
model: "gemini-3.7-flash",
input: largePrompt,
});
console.log(
interaction.usage
);
ตรวจ
total_cached_tokens
จาก Usage Object
ควรเก็บ Metric นี้ใน Production
สมมติมี
Total Input Tokens
=
1,000,000
และ
Cached Tokens
=
700,000
Cache Hit Ratio โดยประมาณคือ
700,000
÷
1,000,000
=
70%
สามารถเก็บ Metric
cache_hit_tokens
total_input_tokens
cache_hit_ratio
เพื่อดูว่า Prompt Architecture มีประสิทธิภาพหรือไม่
Interactions API รองรับ Stateful Conversation ผ่าน
previous_interaction_id
Google ระบุว่าการใช้ Stateful Conversation ทำให้ระบบใช้ Implicit Caching กับ Conversation History ได้ง่ายขึ้น
ตัวอย่าง
first = client.interactions.create(
model="gemini-3.7-flash",
input="อ่านข้อความยาวนี้...",
)
second = client.interactions.create(
model="gemini-3.7-flash",
previous_interaction_id=first.id,
input="สรุปประเด็นที่ 2 เพิ่ม",
)
เหมาะกับ Chat หรือ Agent ที่มี Context ต่อเนื่อง
ใช้
previous_interaction_id
ระบบจัดการ Conversation State ให้
และ Google ระบุว่าเหมาะกับการใช้ Implicit Cache ของ Conversation History
Application ส่ง History เอง
Implicit Caching ก็ยังรองรับ
แต่ Developer ต้องรักษา Prefix ให้เสถียรเองมากขึ้น
นี่คือข้อมูลสำคัญที่สุดของบทความนี้
ถ้าใช้ Interactions API
client.interactions.create()
ปัจจุบันรองรับ
Implicit Caching
เท่านั้น
ไม่รองรับการสร้างและอ้างอิง Cache Object ด้วยตัวเอง
หากต้องการ
Create Cache
Set TTL
Reuse Cache ID
Update TTL
Delete Cache
ต้องใช้
generateContent API
สำหรับ Explicit Caching
Explicit Caching คือ Developer สร้าง Cache Object โดยตรง
ตัวอย่าง
Large PDF
↓
Create Cached Content
↓
cache.name
↓
Question 1
Question 2
Question 3
Developer สามารถกำหนด
TTL
และควบคุม Lifecycle ของ Cache เอง
ทำให้ Cost Saving Predictable มากกว่า Implicit Caching สำหรับ Workload ที่ Context เดิมถูกใช้ซ้ำจำนวนมาก
Google ยก Use Case เช่น
Large System Prompt
+
หลาย User Requests
Video
↓
ถาม 100 คำถาม
PDF Set
↓
Recurring Queries
Repository
↓
Explain
Debug
Refactor
Security Review
นี่เป็น Workload ที่ Context ซ้ำชัดเจน
สำหรับ Explicit Caching ใช้ GenerateContent API
Python SDK ยังคงเป็น
google-genai
ตัวอย่าง
from google import genai
from google.genai import types
client = genai.Client()
document = client.files.upload(
file="manual.pdf",
config={
"mime_type": "application/pdf",
},
)
cache = client.caches.create(
model="gemini-3.7-flash",
config=types.CreateCachedContentConfig(
system_instruction=(
"ตอบคำถามโดยใช้ข้อมูล"
"จากเอกสารนี้เท่านั้น"
),
contents=[
document,
],
ttl="3600s",
),
)
response = client.models.generate_content(
model="gemini-3.7-flash",
contents=(
"สรุปเงื่อนไขการรับประกัน"
),
config=types.GenerateContentConfig(
cached_content=cache.name,
),
)
print(response.text)
print(response.usage_metadata)
จุดที่ต้องสังเกตคือ Explicit Cache ใช้
client.caches.create()
และการ Generate ใช้
client.models.generate_content()
ไม่ใช่ client.interactions.create()
document = client.files.upload(...)
cache = client.caches.create(...)
gemini-3.7-flash
3600s
เท่ากับประมาณ 1 ชั่วโมง
cached_content=cache.name
จากนั้น Prompt ใหม่ไม่ต้องประกอบ Document เดิมเข้า Request ทุกครั้งในรูปแบบเดิม
TTL ย่อมาจาก
Time To Live
หมายถึงระยะเวลาที่ Explicit Cache ถูกเก็บไว้
ตัวอย่าง
300s
=
5 นาที
3600s
=
1 ชั่วโมง
7200s
=
2 ชั่วโมง
เมื่อหมด TTL Cache จะ Expire
Explicit Cache ไม่ใช่ Storage ฟรี
Google คิดค่าใช้จ่ายตาม
Cached Token Count
×
Storage Duration
ดังนั้น Cache
1M tokens
เก็บ
1 ชั่วโมง
ถูกกว่าเก็บ
24 ชั่วโมง
หากใช้งานจริงเพียง 20 นาที
การตั้ง TTL 7 วันอาจไม่คุ้ม
ใช้หลัก
Cache จะถูกถามซ้ำนานแค่ไหน?
ตัวอย่าง
30–60 นาที
10–30 นาที
1–3 ชั่วโมง
อาจยาวขึ้นตาม Requirement
ควรคำนวณ Cost ก่อนตั้ง TTL ยาวมาก
Caching คุ้มเมื่อ
เงินที่ประหยัดจาก Cached Input
>
ค่า Cache Storage
สมมติ Cache ใหญ่ แต่ใช้ซ้ำเพียงครั้งเดียว
อาจไม่คุ้ม
แต่ถ้าใช้ซ้ำ
10
100
1,000
ครั้ง โอกาสคุ้มจะสูงขึ้น
ดังนั้นควรดู
Reuse Count
ด้วย
สมมติ
C = Cached Tokens
R = จำนวนครั้งที่ Reuse
Pi = ราคา Input ปกติ
Pc = ราคา Cached Input
Ps = Storage Price
T = ชั่วโมงที่เก็บ
ต้นทุนไม่มี Cacheโดยประมาณ
C × R × Pi
ต้นทุน Cache
C × R × Pc
+
C × T × Ps
โดยต้องปรับหน่วยต่อ 1M Tokens ตาม Pricing
สูตรนี้ช่วยประมาณได้ว่าควร Cache หรือไม่
สำหรับ Gemini 3.7 Flash ถึงสิ้นปี 2026
Normal Input
=
$0.75 / 1M
Cached Input
=
$0.075 / 1M
Storage
=
$0.50 / 1M tokens / hour
ถ้า Context 1M Tokens ใช้ซ้ำ 10 ครั้งภายใน 1 ชั่วโมง
ไม่ Cache:
10 × $0.75
=
$7.50
ใช้ Explicit Cache โดยประมาณ:
10 × $0.075
+
$0.50 storage
=
$1.25
ยังไม่รวม Output และ Input ใหม่
ใน Workload ลักษณะนี้ Cache มีโอกาสคุ้มมาก
Context 1M Tokens
Reuse เพียง 1 ครั้ง
Cached Usage:
$0.075
Storage 1 ชั่วโมง:
$0.50
รวมส่วน Cache ประมาณ
$0.575
เทียบ Input ปกติ
$0.75
ยังอาจประหยัด แต่ Margin ต่างกันไม่มากเท่ากรณี Reuse จำนวนมาก และต้องคำนึงถึง Complexity ในการสร้าง Cache
ดังนั้น Explicit Cache เหมาะที่สุดกับ Repeated Workload
SDK ปัจจุบันคือ
@google/genai
ตัวอย่างแนวคิด
import {
GoogleGenAI,
createUserContent,
createPartFromUri,
} from "@google/genai";
const ai = new GoogleGenAI({});
const doc = await ai.files.upload({
file: "manual.txt",
config: {
mimeType: "text/plain",
},
});
const cache = await ai.caches.create({
model: "gemini-3.7-flash",
config: {
contents: createUserContent(
createPartFromUri(
doc.uri,
doc.mimeType
)
),
systemInstruction:
"ตอบจากเอกสารนี้เท่านั้น",
ttl: "3600s",
},
});
const response =
await ai.models.generateContent({
model: "gemini-3.7-flash",
contents:
"สรุปข้อกำหนดสำคัญ 5 ข้อ",
config: {
cachedContent: cache.name,
},
});
console.log(response.text);
หลักการเหมือน Python
ได้
Explicit Cache สามารถเปลี่ยน
ttl
หรือ
expire_time
ได้
แต่ Google ระบุว่าไม่รองรับการแก้ Content ส่วนอื่นของ Cache
ดังนั้นถ้า Document เปลี่ยน
มักต้องสร้าง Cache ใหม่
ตัวอย่าง
from google.genai import types
client.caches.update(
name=cache.name,
config=types.UpdateCachedContentConfig(
ttl="7200s"
),
)
เปลี่ยน TTL เป็น
2 ชั่วโมง
โดยประมาณ
ได้
Python
client.caches.delete(
cache.name
)
JavaScript
await ai.caches.delete({
name: cache.name,
});
ถ้ารู้ว่า Workflow จบแล้ว ไม่จำเป็นต้องรอ TTL หมด
การลบ Cache ที่ไม่ใช้ช่วยลด Storage Duration ที่ไม่จำเป็น
สามารถเรียกดู Cache ที่มีอยู่
เหมาะกับ
Application ขนาดใหญ่ควรเก็บ Mapping เช่น
document_id
↓
cache_name
↓
expire_time
ไม่ควรสร้าง Cache ใหม่ทุก Request โดยไม่ตรวจว่ามีของเดิมอยู่หรือไม่
ตัวอย่างไม่ดี
Question 1
↓
Create Cache
↓
Ask
↓
Done
Question 2
↓
Create Cache ใหม่
↓
Ask
แบบนี้เสียประโยชน์หลักของ Caching
ควรเป็น
Create Cache ครั้งเดียว
↓
Question 1
Question 2
Question 3
Question 4
ภายใน Lifecycle ที่เหมาะสม
บาง Application มี System Prompt เช่น
20,000 tokens
ประกอบด้วย
ถ้า Prompt ส่วนนี้เหมือนทุก Request จะเป็น Candidate ที่ดีสำหรับ Caching
โดยเฉพาะเมื่อ Traffic สูง
ตัวอย่าง
คู่มือสินค้า 400 หน้า
User ถาม
คำถาม 1
คำถาม 2
คำถาม 3
...
ถ้า Context เดิมถูกใช้ซ้ำจำนวนมาก Caching สามารถลดค่า Input ที่ซ้ำได้อย่างมีนัยสำคัญ
Video สามารถสร้าง Context จำนวนมาก
ถ้าต้องถามวิดีโอเดิมหลายคำถาม
Video
↓
Cache
↓
หาเหตุการณ์ช่วงต้น
↓
สรุปตัวละคร
↓
หา Timestamp
↓
วิเคราะห์ฉาก
ดีกว่าประมวลผล Video Context ใหม่เต็มรูปแบบทุกคำถาม
ตัวอย่าง
Repository
↓
Cache
แล้วถาม
Architecture คืออะไร
Bug อยู่ตรงไหน
Refactor Module นี้
ตรวจ Security
หาก Codebase เดิมเป็น Common Context การ Reuse สามารถช่วยเรื่อง Cost ได้
หาก Cached Document Version 1
manual-v1.pdf
ถูกแก้เป็น
manual-v2.pdf
อย่าใช้ Cache เดิมต่อโดยไม่คิด
เพราะโมเดลจะอ้าง Context เก่า
ควร
สร้าง Cache ใหม่
↓
Switch Traffic
↓
Delete Cache เก่า
นี่คือ Cache Invalidation
Cache ลด Cost แต่เพิ่มปัญหาใหม่คือ
ข้อมูลใน Cache ยังใหม่หรือไม่?
ตัวอย่าง
ข้อมูลที่เปลี่ยนบ่อยไม่เหมาะกับ Explicit Cache TTL ยาว
โครงสร้างที่ดีคือ
Static Context
→ Cache
Dynamic Context
→ Request ปัจจุบัน
ตัวอย่าง
Product Manual
→ Cached
Current Inventory
→ Function Call / Dynamic Input
ไม่ควร Cache Inventory 24 ชั่วโมงถ้าสต็อกเปลี่ยนทุกนาที
อย่าใช้ Context Cache เป็น Database ของ Application
Cache ถูกออกแบบเพื่อ
ลดการประมวลผล Context ซ้ำ
ไม่ใช่เก็บ Source of Truth
Source จริงควรอยู่ใน
ตามประเภทข้อมูล
Google ระบุว่าโมเดลไม่ได้แยก Cached Tokens ออกจาก Input Tokens ในเชิงความหมาย
Cached Content ทำหน้าที่เป็น
Prefix
ของ Prompt
ดังนั้น Context Window ยังนับ Cached Content อยู่
Caching ลดราคา แต่ไม่ได้เพิ่ม Context Window อย่างไม่มีขีดจำกัด
สมมติ Model รองรับ Context สูงสุดระดับหนึ่ง
การ Cache ข้อมูล
1M tokens
ไม่ได้ทำให้สามารถเพิ่มอีก
1M
+
1M
+
1M
โดยไม่สน Context Limit
Cached Tokens ยังอยู่ภายใต้ Token/Context Limit ของ Model
สำหรับ Explicit Caching Google ระบุว่าไม่มี Rate/Usage Limit พิเศษที่ทำให้หลบข้อจำกัด GenerateContent
Standard Rate Limits ยังคงมีผล
ดังนั้น Cache ไม่ใช่วิธีแก้
429
ทุกกรณี
หากปัญหาเป็น RPM ยังต้องจัดการ Request Rate
ต้องระวังการตีความ
แม้ Cached Input มีราคาต่ำกว่า แต่ Cached Tokens ยังเป็นส่วนหนึ่งของ Context และข้อจำกัด Token ที่เกี่ยวข้องตาม API/Model
ดังนั้นอย่าออกแบบว่า
Cache
=
ไม่ใช้ Token
ไม่ถูกต้อง
Caching เป็น Cost Optimization เป็นหลัก
นี่เป็นจุดที่ต่างจาก Implicit Cache
Implicit
Google จัดการให้
Explicit
Developer สร้าง Cache
+
กำหนด TTL
+
เสีย Storage ตามระยะเวลา
ดังนั้น Explicit Cache ต้องมี Lifecycle Management
ระบบจริงควรมี
Cache Created
↓
Used
↓
Workflow Completed
↓
Delete
หรือปล่อยให้ Expire ตาม TTL ที่เหมาะสม
ไม่ควรตั้ง TTL ยาวสุดโดยไม่จำเป็น
Database อาจมี
cache_name
document_id
model
created_at
expire_at
token_count
status
ช่วยให้ Application รู้ว่า
หาก Context มี
ต้องตรวจ Data Policy และ Security Requirement ของระบบก่อน Cache
อย่าคิดว่าเพราะเป็น Cache แล้ว Data Security ไม่สำคัญ
ใช้ Data Minimization เช่นเดิม
ควรเก็บ
input_tokens
cached_tokens
cache_hit_ratio
storage_hours
cache_create_count
จากนั้นวิเคราะห์
ต้นทุนก่อน Cache
vs
ต้นทุนหลัง Cache
หาก Cache Hit ต่ำมาก Explicit Caching อาจไม่คุ้ม Complexity
อาจแสดง
Total Input Tokens
100M
Cached Tokens
75M
Cache Hit Ratio
75%
Cache Storage Cost
$X
Estimated Savings
$Y
ช่วยให้ตัดสินได้จากข้อมูลจริง
Group A
No explicit cache
Group B
Explicit cache
เปรียบเทียบ
โดยใช้ Workload ใกล้เคียงกัน
ถ้า Saving ต่ำมากอาจใช้ Implicit Caching อย่างเดียวก็พอ
Caching ถูกออกแบบเพื่อเพิ่มประสิทธิภาพและลด Cost แต่ไม่ควรรับประกันว่า Latency ทุก Request จะลดตามสัดส่วนเดียวกับราคา
ต้อง Benchmark
p50
p95
p99
ใน Application จริง
ประโยชน์หลักที่ควรวัดอย่างชัดเจนคือ
Cached Tokens
และ
Cost
| หัวข้อ | Implicit | Explicit |
|---|---|---|
| เปิดอัตโนมัติ | ใช่ | ไม่ |
| Developer สร้าง Cache | ไม่ | ใช่ |
| กำหนด TTL | ไม่ | ใช่ |
| รับประกัน Cache Hit | ไม่ | ควบคุมการ Reuse ได้มากกว่า |
| Interactions API | รองรับ | ไม่รองรับ |
| generateContent | รองรับ | รองรับ |
| Storage Cost ที่จัดการเอง | ไม่ | มี |
| Complexity | ต่ำ | สูงกว่า |
| เหมาะกับ | Application ทั่วไป | Context ใหญ่ที่ Reuse ชัดเจน |
สำหรับ Project ใหม่ควรเริ่มจาก Implicit ก่อน
ใช้ Decision Tree
ใช้ Interactions API?
↓
ใช่
→ Implicit Caching
ไม่
↓
Context ใหญ่และใช้ซ้ำมาก?
↓
ไม่
→ Implicit เพียงพอ
ใช่
↓
ต้องการควบคุม TTL/Reuse?
↓
ใช่
→ พิจารณา Explicit Cache
อย่าเปลี่ยนจาก Interactions API ไป GenerateContent เพียงเพราะเห็นคำว่า Cache ถ้า Saving ไม่คุ้มกับ Architecture ที่เพิ่มขึ้น
Google ปัจจุบันให้ Interactions API เป็นเส้นทางหลักสำหรับการสร้าง Application ใหม่
ดังนั้น Workflow ที่เหมาะสำหรับหลาย Project คือ
Interactions API
↓
Implicit Caching
↓
วัด total_cached_tokens
↓
Optimize Prompt Prefix
ก่อน
ถ้าพบว่า Workload มี Context ใหญ่ที่ใช้ซ้ำมากและ Explicit Cache จะสร้าง Saving ชัดเจน จึงค่อยพิจารณา GenerateContent API สำหรับส่วนงานนั้น
ในทาง Architecture สามารถแยก Feature ได้
เช่น
Chat / Agent
→ Interactions API
Large repeated document analysis
→ generateContent + Explicit Cache
ไม่จำเป็นต้องบังคับทั้งระบบใช้ API แบบเดียวถ้า Requirement ต่างกัน
แต่ควรเพิ่ม Abstraction Layer เพื่อไม่ให้ Business Logic ผูกกับ API Structure มากเกินไป
Architecture
Application
↓
AI Service
├── InteractionService
└── CachedDocumentService
ทำให้ในอนาคตเปลี่ยน
ได้ง่ายกว่าให้ทุก Controller เรียก SDK โดยตรง
มี
Manual
50,000 tokens
User ถามวันละ
10,000 Questions
ข้อมูล Manual เปลี่ยนเดือนละครั้ง
นี่เป็น Candidate ที่ดีมากสำหรับ Caching
สามารถออกแบบ
Manual Version
↓
Cache / Stable Prefix
↓
User Question
เมื่อ Manual Update
Create new cache
↓
Switch version
↓
Delete old cache
มีวิดีโอ 1 ชั่วโมง
ต้องการถาม
แทนประมวลผล Video Context ซ้ำในทุก Request Explicit Cache อาจมีประโยชน์สูง
โดยเฉพาะ Session ที่มีคำถามต่อเนื่องหลายครั้ง
มี Repository ขนาดใหญ่
Developer ถาม
Architecture
Bug
Refactor
Unit Test
Security
Common Code Context สามารถเป็น Candidate สำหรับ Cache
แต่ Repository เปลี่ยนหลัง Commit ใหม่
จึงต้องมี
commit_hash
หรือ Version เป็นส่วนหนึ่งของ Cache Key/Metadata
แม้ Gemini จะให้ Cache Name เอง แต่ Application ควรมี Logical Key เช่น
document_id + document_version + model
ตัวอย่าง
manual-42:v3:gemini-3.7-flash
ช่วยป้องกันการใช้ Cache ผิด Version
ถ้า Application มีข้อมูลเฉพาะ User
ต้องป้องกันกรณี
User A Cache
→ User B Request
อาจทำให้ข้อมูลรั่ว
Cache Mapping ต้องรวม Tenant/User Scope เมื่อข้อมูลไม่ใช่ Shared Public Context
นี่เป็น Security Requirement สำคัญ
Key อาจเป็น
tenant_id
+
document_id
+
version
+
model
ก่อนใช้ Cache ทุกครั้งต้องตรวจว่า Request มีสิทธิ์เข้าถึง Context นั้นจริง
AI Caching ไม่แทน Authorization
System Prompt, PDF, Video, Code
อย่าเดาจาก File Size
ให้ส่วนซ้ำอยู่ด้านหน้า
ช่วย Conversation History Caching
total_cached_tokensดู Cache Hit จริง
ไม่เพิ่ม Complexity โดยไม่มีเหตุผล
ลด Storage Cost
ป้องกันข้อมูลเก่า
ดู Saving จริง
ปัจจุบันไม่รองรับ
ต้องใช้ GenerateContent สำหรับ Explicit Cache
เปิดอัตโนมัติใน Model ที่รองรับ
ไม่รับประกัน
โอกาส/สิทธิ์ Cache ตาม Minimum ไม่เข้าเงื่อนไข
ลด Cache Hit
ทำ Common Prefix เปลี่ยน
เสียประโยชน์การ Reuse
Storage Cost เพิ่ม
เสีย Storage โดยไม่จำเป็น
ข้อมูลล้าสมัย
ไม่จริง
ไม่จริง
ไม่รู้ว่าประหยัดจริงหรือไม่
เสี่ยงข้อมูลรั่ว
ถ้าเป็น Interactions API ให้เริ่มจาก Implicit
เช่น
gemini-3.7-flash
3.7 Flash ปัจจุบัน 4,096 Tokens
สิ่งที่ซ้ำหลาย Request
ไม่ต้องเปิด Implicit Cache
ดู total_cached_tokens
เปรียบเทียบ Cost
พิจารณา Explicit Cache
สำหรับ Explicit Caching
ให้ตรง Session
ป้องกันข้อมูลเก่า
สำหรับระบบของ comsiam จุดเริ่มต้นที่เหมาะคือใช้ Interactions API และจัด Prompt Prefix ให้คงที่ก่อน แล้ววัด total_cached_tokens จริง หากพบ Context ใหญ่ที่ถูกถามซ้ำจำนวนมากจึงค่อยเพิ่ม Explicit Caching เฉพาะ Feature นั้น
requests
input_tokens
cached_tokens
output_tokens
cache_hit_ratio
cache_storage_hours
cost_without_cache_estimate
actual_cost
estimated_savings
ถ้าใช้ Explicit Cache เพิ่ม
cache_create_count
cache_delete_count
expired_cache_count
ด้วย
ไม่ควร
User ID: 123
Current Time: 10:55
Random Request ID: ABC
[Large Document]
Question
เพราะ Prefix เปลี่ยนทุกครั้ง
ควรเป็น
[Stable System Instruction]
[Large Shared Document]
[Dynamic User Information]
[Question]
เมื่อ Business Logic อนุญาต
นี่เป็น Optimization ที่ทำได้โดยไม่ต้องเปลี่ยน API
ใช้
SYSTEM_PROMPT_V1
ให้คงที่ใน Request ชุดหนึ่ง
เมื่อพร้อม Deploy
SYSTEM_PROMPT_V2
จึง Switch
อย่าสร้างข้อความ System Prompt ที่แตกต่างเล็กน้อยแบบสุ่มทุก Request เพราะลด Prefix Reuse
Caching เหมาะกับ
Context เดิม
ใช้ซ้ำ
RAG เหมาะกับ
Dataset ใหญ่มาก
ค้นเฉพาะส่วนที่เกี่ยว
สมมติ Knowledge Base มี 10 ล้าน Documents
ไม่ควร Cache ทั้งหมดเข้า Context
ควร
Retrieve relevant documents
↓
Send only relevant context
และอาจ Cache ส่วนที่ถูกใช้ซ้ำภายหลัง
Caching กับ RAG สามารถใช้ร่วมกันได้
เก็บ Source of Truth
หา Context ที่เกี่ยว
ลดต้นทุน Context ที่ถูกใช้ซ้ำ
ทั้งสามทำหน้าที่ต่างกัน
Architecture ที่ดีอาจเป็น
Database / Documents
↓
Retrieval
↓
Common Context
↓
Cache
↓
Gemini
ตาม Use Case
Pricing แตกต่างตาม Model และ Processing Mode
บางรายการใน Pricing แสดง Context Caching Free หรือ Not Available แตกต่างกันตาม Mode/Model
ดังนั้นอย่าตั้ง Rule ว่า
Free Tier
=
Caching ฟรีทุก Model
หรือ
Paid เท่านั้นทุกกรณี
ควรตรวจ Pricing ของ Model ที่กำลังใช้โดยตรง
โดยเฉพาะ Gemini 3.7 Flash ปัจจุบันมีราคาช่วงถึง
31 ธันวาคม 2026
และราคาใหม่เริ่ม
1 มกราคม 2027
ดังนั้นระบบที่คำนวณ Budget ปีหน้าต้องใช้ราคาใหม่ด้วย
ไม่ควร Hard-code ราคาปัจจุบันใน Business Logic ถาวร
คือกลไกที่ช่วยให้ Gemini ใช้ Input Context เดิมซ้ำในราคาที่ต่ำกว่าการประมวลผล Input ปกติเมื่อ Cache Hit เหมาะกับ System Prompt, Document, Video หรือ Code ขนาดใหญ่ที่ถูกใช้กับหลาย Request
ได้ แต่ปัจจุบันรองรับ Implicit Caching เท่านั้น ซึ่งเปิดอัตโนมัติ ไม่รองรับการสร้าง Explicit Cache Object ด้วยตัวเอง
ไม่ต้องเปิด Gemini 2.5 และรุ่นใหม่กว่าที่รองรับเปิด Implicit Caching ให้อัตโนมัติ Developer เพียงจัด Common Prefix ให้เสถียรและตรวจ Cached Tokens จาก Usage
Minimum Input Token สำหรับ Context Caching ปัจจุบันคือ 4,096 Tokens
สำหรับ Interactions API สามารถดูจำนวน Cached Tokens จาก usage.total_cached_tokens
ไม่ได้ในปัจจุบัน หากต้องการสร้าง Cache Object, ตั้ง TTL และส่ง cached_content เอง ต้องใช้ GenerateContent API
Context Caching เป็นหนึ่งในวิธีลดค่าใช้จ่าย Gemini API ที่มีประโยชน์มากที่สุดเมื่อ Application มี Context ขนาดใหญ่ที่ถูกส่งซ้ำหลายครั้ง เช่น System Instruction, PDF, Video, Code Repository หรือ Conversation History
Gemini API ปัจจุบันมีทั้ง Implicit และ Explicit Caching แต่สำหรับ Interactions API รองรับเฉพาะ Implicit Caching และเปิดอัตโนมัติสำหรับ Gemini 2.5 และรุ่นใหม่กว่าที่รองรับ Developer ไม่จำเป็นต้องสร้าง Cache หรือกำหนด TTL เอง
สำหรับ Gemini 3.7 Flash ต้องมี Input อย่างน้อย 4,096 Tokens ตาม Minimum Caching Threshold ปัจจุบัน และสามารถเพิ่มโอกาส Cache Hit โดยวาง Context ใหญ่ที่เหมือนกันไว้ด้านหน้า Prompt พร้อมรักษา Prefix ให้คงที่ใน Request ที่เกิดใกล้กัน
Cache Hit สามารถตรวจได้จาก usage.total_cached_tokens จึงควร Monitor ตัวเลขจริงแทนการเดาว่า Cache ทำงานหรือไม่
หากต้องการควบคุม Cache เองแบบ Explicit เช่น สร้าง Cache Object, กำหนด TTL, Update Expiry และ Delete Cache ต้องใช้ GenerateContent API โดยต้นทุนประกอบด้วย Cached Token Usage และ Storage Duration
Explicit Cache เหมาะกับงานที่ Context เดิมถูกใช้ซ้ำจำนวนมาก เช่นถาม PDF เดิมหลายร้อยครั้ง วิเคราะห์วิดีโอเดิมซ้ำ หรือทำ Code Review กับ Repository เดิม แต่ไม่ควรใช้กับข้อมูลที่เปลี่ยนบ่อยโดยไม่มี Cache Invalidation Strategy
สำหรับ comsiam แนวทางที่คุ้มที่สุดคือเริ่มจาก Implicit Caching ที่มากับ Interactions API ก่อน จัด Prompt ให้ Common Prefix เสถียร เก็บ Cached Token Metrics และวัด Cost จริง จากนั้นใช้ Explicit Caching เฉพาะ Workflow ที่พิสูจน์แล้วว่ามี Reuse สูงเพียงพอที่จะคุ้มกับ Storage Cost และ Complexity ที่เพิ่มขึ้น