Insights · Applied AI
Cutting LLM cost with a Thai decision model: where it worked, where it failed
A lot of what we ask large language models to do is not writing at all. It is picking one answer from a list we already know. We tested whether a small, self-hosted Thai model could take that work off the LLM, and published the misses along with the wins.
The problem: paying LLM prices for a multiple-choice question
Look inside most AI assistants and you will find steps like these: which topic is this message about, which team should handle it, is this urgent, does this document answer the question. Each one is a multiple-choice decision. Yet the usual way to build them is to send the whole conversation to a general-purpose LLM, ask for JSON back, and parse the reply, which occasionally breaks.
That works, but you pay for every input token, wait for text generation, and get no reliable signal for how sure the model is.
What a “decision model” is
In September 2026 TypeSafe AI released Jev, a model that does not generate text. You give it a state (text or JSON) and typed questions, and it returns a typed answer with probabilities and a confidence value. It offers three question types:
- Choice: pick one of up to 255 options (intent, routing target).
- Score: rate against 2–10 ordered levels (difficulty, urgency).
- Noul: a yes/no statement returned as a probability from 0 to 1 (“does this passage answer the question?”).
An open re-implementation, OpenThai-SystemOne by iApp, uses the same request contract. It is Apache-2.0, has 0.8B parameters, and received roughly 5B tokens of extra Thai training. Because it can run on your own infrastructure, it was the one we tested.
What we tested
Both tests ran offline on a development machine (CPU, no GPU), in September 2026.
Test A: routing website visitors · 35 messages, mixed Thai and English
The job: sort messages of the kind visitors send to VIRA, the assistant on this website, into a fixed set of topics, including “other / off-topic”. We wrote the 35 test messages ourselves; no real visitor conversations were used.
35 hand-written messages. A small test set: read it as a direction, not a guarantee.
The most useful finding was the confidence value. On its own, the model caught only 1 of 5 off-topic messages. But those messages all came back with very low confidence (0.04–0.21). Treating anything under 0.5 as “other” caught all five. If you use a model like this, the confidence score is the product, not a detail.
A hybrid setup (keywords first, model only when keywords say “other”) reached the same 28/35 while calling the model on 21 of 35 messages. The remaining errors were in categories whose meanings overlap. Sharper category descriptions are the fix.
Test B: does this document answer the question? · 8 short sample HR policy passages we wrote, 16 questions
| Task | Result | Verdict |
|---|---|---|
| Judge each passage relevant / not relevant | Kept 8/10 relevant · dropped 13/16 irrelevant | Good enough to filter context before the LLM |
| Pick the document that answers (answer exists) | 9/10 | Good |
| Say “no document answers this” | 1/6 | Weak. A question about personal leave scored 0.90 against the annual-leave policy. |
Use it to trim context, never to decide “we have no information on this”. In our tests it was confidently wrong exactly where a wrong answer hurts most.
What it does to cost
Take one routing call. Assumptions: a typical LLM routing prompt carries about 1,500 input tokens (instructions plus chat history) and the reply is about 80. The decision model needs only the latest message and two previous turns, about 400 tokens.
| Option | Price basis | Per routing call |
|---|---|---|
| General LLM, 70B class | $0.59 in / $0.79 out per 1M tokens | ≈ $0.00095 |
| Jev (hosted) | $0.042 per 1M input · output free | ≈ $0.000017 |
| OpenThai-SystemOne (self-hosted) | No per-call fee · ~1.4 GB RAM in bf16 | Your own compute |
On these assumptions, the hosted decision model is roughly 56× cheaper per decision (we did not test the hosted service; our tests used the self-hosted model). Our assumption is that 20–30% of messages (low confidence, or genuinely open-ended) would still go to the LLM, so the saving applies to the rest. The same “decide cheaply first” idea also applies one level up: score how hard a request is, and send the easy ones to a cheaper model.
The pattern we recommend
Log the model version, the choice and the confidence for every decision. That is how you tune the threshold and how you answer an auditor later.
Before you put this in production
- Personal data (PDPA). According to public reports (September 2026), TypeSafe does not use requests for training and offers zero data retention to enterprise customers, but does not state where data is processed. Check the vendor's current terms yourself. A self-hosted model can keep personal data on infrastructure you control, but you still need a lawful basis, access control and a retention rule. This is engineering guidance, not legal advice.
- Never expose the model server directly. The reference server ships without authentication. Keep it on an internal network or behind your own gateway.
- User text can steer the answer. A crafted message can change a classification. Do not make it the only gate for anything security-related; pair it with rules you control.
- Pin the model version. A “latest” tag can change underneath you and silently move your confidence threshold.
- Know the blind spots. Counting, date comparison, multi-step reasoning, and saying “there is no answer”.
The takeaway
A decision model does not replace your LLM. It is a cheap first step in front of it. It is useful when the answer is a choice, when you actually use the confidence value, and when the data stays where it should. Where the job is explaining, writing or reasoning, keep the LLM.
If you want to know where a decision step like this would pay back in your workflows, tell us about the process, or ask VIRA in the corner of this page.
Sources
- Jev (AI model), Wikipedia
- A deep dive into Jev, flaviocopes.com (pricing, limits, weaknesses)
- iapp/OpenThai-SystemOne, Hugging Face (licence, size, Thai training)
- 70B-class LLM price: one public list price for Llama 3.3 70B inference, September 2026. Prices vary by provider.
Insights · Applied AI
ลดต้นทุน LLM ด้วยโมเดลตัดสินใจภาษาไทย: ได้ผลตรงไหน พลาดตรงไหน
งานที่เราสั่งให้ LLM ทำจำนวนมากไม่ใช่การเขียนข้อความ แต่คือการเลือกคำตอบหนึ่งข้อจากตัวเลือกที่รู้อยู่แล้ว เราทดสอบว่าโมเดลภาษาไทยขนาดเล็กที่รันบนระบบของเราเองจะรับงานส่วนนี้แทน LLM ได้แค่ไหน และเผยแพร่ทั้งส่วนที่ได้ผลและส่วนที่พลาด
ปัญหา: จ่ายราคา LLM เพื่อตอบข้อสอบปรนัย
ลองเปิดดูข้างใน AI assistant ส่วนใหญ่ จะเจอขั้นตอนแบบนี้ เช่น ข้อความนี้เป็นเรื่องอะไร ควรส่งให้ทีมไหน เร่งด่วนไหม เอกสารนี้ตอบคำถามได้หรือเปล่า ทุกข้อคือการเลือกจากตัวเลือก แต่วิธีที่นิยมทำกันคือส่งบทสนทนาทั้งหมดให้ LLM ทั่วไป ขอให้ตอบเป็น JSON แล้วแกะคำตอบ ซึ่งบางครั้งก็แกะไม่ออก
วิธีนี้ใช้งานได้ แต่ต้องจ่ายค่า input token ทุกตัว ต้องรอโมเดลสร้างข้อความ และไม่มีตัวเลขที่เชื่อถือได้ว่าโมเดลมั่นใจแค่ไหน
“โมเดลตัดสินใจ” คืออะไร
เดือนกันยายน 2026 TypeSafe AI เปิดตัว Jev ซึ่งเป็นโมเดลที่ไม่สร้างข้อความ เราส่ง “สถานะ” (ข้อความหรือ JSON) กับคำถามที่กำหนดชนิดไว้ แล้วโมเดลคืนคำตอบที่มีชนิดชัดเจน พร้อมความน่าจะเป็นและค่า confidence มีคำถาม 3 แบบ:
- Choice เลือก 1 จากตัวเลือกได้สูงสุด 255 ข้อ เช่น หัวข้อของข้อความ หรือปลายทางที่จะส่งต่อ
- Score ให้คะแนนตามระดับที่เรียงไว้ 2–10 ระดับ เช่น ความยาก หรือความเร่งด่วน
- Noul ตอบใช่/ไม่ใช่เป็นความน่าจะเป็น 0–1 เช่น “ข้อความนี้ตอบคำถามได้ไหม”
ฝั่งโมเดลเปิดมี OpenThai-SystemOne ของ iApp ที่ใช้รูปแบบ API เดียวกัน เป็นสัญญาอนุญาต Apache-2.0 ขนาด 0.8B parameters และฝึกภาษาไทยเพิ่มราว 5B token เราเลือกตัวนี้มาทดสอบเพราะรันบนระบบของตัวเองได้
เราทดสอบอะไร
ทั้งสองชุดรันแบบ offline บนเครื่องพัฒนา (CPU ไม่ใช้ GPU) ในเดือนกันยายน 2026
ชุด ก. แยกหัวข้อข้อความผู้เข้าชมเว็บ · 35 ข้อความ ไทยปนอังกฤษ
โจทย์คือจัดข้อความแบบที่ผู้เข้าชมมักพิมพ์คุยกับ VIRA ผู้ช่วยบนเว็บนี้ ลงหัวข้อที่กำหนดไว้ รวมถึงหัวข้อ “อื่นๆ / นอกเรื่อง” ข้อความทดสอบ 35 ข้อเราเขียนขึ้นเอง ไม่ได้ใช้บทสนทนาจริงของผู้เข้าชม
ชุดทดสอบ 35 ข้อความที่เขียนขึ้นเอง ถือเป็นแนวโน้ม ไม่ใช่การรับประกันผล
สิ่งที่มีประโยชน์ที่สุดคือค่า confidence ถ้าใช้โมเดลเฉยๆ จะจับข้อความนอกเรื่องได้แค่ 1 ใน 5 แต่ข้อความพวกนี้ได้ confidence ต่ำมากทุกข้อ (0.04–0.21) พอตั้งเกณฑ์ว่าต่ำกว่า 0.5 ให้ถือเป็น “อื่นๆ” ก็จับได้ครบทั้ง 5 ข้อ ถ้าจะใช้โมเดลแบบนี้ ค่า confidence คือหัวใจ ไม่ใช่รายละเอียดปลีกย่อย
ถ้าใช้แบบผสม คือให้ keyword ตัดสินก่อน แล้วเรียกโมเดลเฉพาะเมื่อ keyword ตอบ “อื่นๆ” ก็ได้ 28/35 เท่ากัน แต่เรียกโมเดลเพียง 21 จาก 35 ครั้ง ข้อที่ยังพลาดอยู่ในหัวข้อที่ความหมายทับกัน ซึ่งแก้ได้ด้วยการเขียนคำอธิบายหัวข้อให้ชัดขึ้น
ชุด ข. เอกสารนี้ตอบคำถามได้ไหม · ข้อความนโยบาย HR ตัวอย่าง 8 ข้อที่เราเขียนขึ้นเอง, 16 คำถาม
| การทดสอบ | ผล | สรุป |
|---|---|---|
| ตัดสินทีละส่วนว่าเกี่ยว / ไม่เกี่ยว | เก็บส่วนที่เกี่ยว 8/10 · ตัดส่วนที่ไม่เกี่ยว 13/16 | ใช้กรอง context ก่อนส่งให้ LLM ได้ |
| เลือกเอกสารที่ตอบได้ (กรณีมีคำตอบจริง) | 9/10 | ดี |
| บอกว่า “ไม่มีเอกสารไหนตอบได้” | 1/6 | อ่อน คำถามเรื่องลากิจได้คะแนน 0.90 ว่านโยบายลาพักร้อนตอบได้ |
ใช้ตัดทอน context ได้ แต่อย่าใช้ตัดสินว่า “ระบบไม่มีข้อมูลเรื่องนี้” ในการทดสอบของเรา โมเดลตอบผิดอย่างมั่นใจตรงจุดที่ผิดแล้วเสียหายที่สุดพอดี
ผลต่อต้นทุน
ลองดูการจัดหัวข้อหนึ่งครั้ง สมมติฐาน: prompt สำหรับจัดหัวข้อด้วย LLM โดยทั่วไปมี input ราว 1,500 token (คำสั่ง + ประวัติแชต) และคำตอบราว 80 token ส่วนโมเดลตัดสินใจต้องการแค่ข้อความล่าสุดกับ 2 turn ก่อนหน้า ราว 400 token
| ทางเลือก | ฐานราคา | ต่อการตัดสินใจ 1 ครั้ง |
|---|---|---|
| LLM ทั่วไประดับ 70B | $0.59 input / $0.79 output ต่อ 1M token | ≈ $0.00095 |
| Jev (บริการบน cloud) | $0.042 ต่อ 1M input · output ไม่คิดเงิน | ≈ $0.000017 |
| OpenThai-SystemOne (รันเอง) | ไม่มีค่าต่อครั้ง · ใช้ RAM ~1.4 GB แบบ bf16 | ต้นทุนเครื่องของเราเอง |
ตามสมมติฐานนี้ โมเดลตัดสินใจแบบบริการบน cloud ถูกลงราว 56 เท่าต่อการตัดสินใจหนึ่งครั้ง (เราไม่ได้ทดสอบบริการบน cloud การทดสอบของเราใช้โมเดลที่รันเอง) และเราสมมติว่าข้อความราว 20–30% (confidence ต่ำ หรือเป็นคำถามปลายเปิดจริงๆ) ยังต้องส่งให้ LLM ส่วนที่ประหยัดได้จึงมาจากข้อความที่เหลือ แนวคิด “ตัดสินใจราคาถูกก่อน” นี้ใช้กับระดับที่สูงขึ้นได้ด้วย คือให้คะแนนความยากของคำขอ แล้วส่งคำขอที่ง่ายไปโมเดลที่ถูกกว่า
รูปแบบที่เราแนะนำ
บันทึกเวอร์ชันโมเดล ผลที่เลือก และค่า confidence ทุกครั้ง ข้อมูลนี้ใช้ปรับเกณฑ์ และใช้ตอบผู้ตรวจสอบได้ในภายหลัง
ก่อนขึ้นระบบจริง
- ข้อมูลส่วนบุคคล (PDPA) จากข้อมูลที่เผยแพร่ ณ กันยายน 2026 TypeSafe ระบุว่าไม่นำคำขอไปฝึกโมเดล และมีตัวเลือกไม่เก็บข้อมูล (zero data retention) สำหรับลูกค้าองค์กร แต่ไม่ได้ระบุว่าประมวลผลที่ภูมิภาคไหน ควรตรวจเงื่อนไขล่าสุดของผู้ให้บริการเอง การรันโมเดลเองช่วยให้เก็บข้อมูลส่วนบุคคลไว้บนระบบที่เราควบคุมได้ แต่ยังต้องมีฐานทางกฎหมาย การควบคุมสิทธิ์เข้าถึง และระยะเวลาเก็บข้อมูล ข้อความนี้เป็นคำแนะนำเชิงวิศวกรรม ไม่ใช่คำแนะนำทางกฎหมาย
- อย่าเปิดเซิร์ฟเวอร์โมเดลออกสู่สาธารณะตรงๆ เซิร์ฟเวอร์ตัวอย่างที่มากับโมเดลไม่มีระบบยืนยันตัวตน ให้วางไว้ในเครือข่ายภายในหรือหลัง gateway ของเราเอง
- ข้อความจากผู้ใช้หลอกให้เปลี่ยนคำตอบได้ ข้อความที่จงใจเขียนมาเปลี่ยนผลการจัดหัวข้อได้ จึงไม่ควรใช้เป็นด่านเดียวของเรื่องความปลอดภัย ให้ใช้คู่กับกฎที่เราควบคุมเอง
- ล็อกเวอร์ชันโมเดล ถ้าใช้ป้าย “latest” โมเดลอาจเปลี่ยนเองโดยไม่รู้ตัว และทำให้เกณฑ์ confidence ที่ตั้งไว้คลาดเคลื่อน
- รู้จุดบอด การนับ การเทียบวันที่ การให้เหตุผลหลายขั้น และการตอบว่า “ไม่มีคำตอบ”
สรุป
โมเดลตัดสินใจไม่ได้มาแทน LLM แต่เป็นด่านแรกราคาถูกที่วางไว้ข้างหน้า มันคุ้มเมื่อคำตอบเป็นการเลือก เมื่อเราใช้ค่า confidence จริงจัง และเมื่อข้อมูลอยู่ในที่ที่ควรอยู่ ส่วนงานที่ต้องอธิบาย เขียน หรือใช้เหตุผล ยังควรให้ LLM ทำ
ถ้าอยากรู้ว่าด่านตัดสินใจแบบนี้จะคุ้มกับงานตรงไหนขององค์กร เล่ากระบวนการให้เราฟัง หรือถาม VIRA ที่มุมหน้าจอได้เลย
แหล่งอ้างอิง
- Jev (AI model), Wikipedia
- A deep dive into Jev, flaviocopes.com (ราคา ข้อจำกัด จุดอ่อน)
- iapp/OpenThai-SystemOne, Hugging Face (สัญญาอนุญาต ขนาด การฝึกภาษาไทย)
- ราคา LLM ระดับ 70B: ราคาประกาศของผู้ให้บริการ inference Llama 3.3 70B รายหนึ่ง ณ กันยายน 2026 ราคาต่างกันไปตามผู้ให้บริการ