Skip to main content

🧠 OpenThaiGPT R1 32b

OpenThaiGPT R1 32b is an advanced 32-billion-parameter Thai language model focused on analytical thinking and reasoning, outperforming larger models such as DeepSeek R1 70b and Typhoon R1 70b despite being less than half their size. The model excels at tasks requiring complex analytical thinking, including mathematics, logic, and coding in the Thai language.

Download the model

OpenThaiGPT R1 32b — Hugging Facehuggingface.co

openthaigpt/openthaigpt-r1-32b-instruct

Highlights

  • State-of-the-art Thai language model that outperforms larger models on mathematics and logical reasoning benchmarks
  • Explicit reasoning capabilities able to show its thought process step by step
  • Significantly smaller size (32b) yet higher performance than 70b models
  • Specialized in analytical thinking in Thai, including complex mathematical and logical problems
  • High coding performance in both Thai and English

Benchmark results

SkyThoughtOpenThaiGPT R1 32bDeepSeek R1 70bTyphoon R1 70b
AIME24-TH56.6733.3353.33
AIME2463.3653.3353.33
MATH500-TH83.875.481
MATH50089.488.8890.2
LiveCodeBench-TH62.1653.1547.75
LiveCodeBench69.6764.9754.79
OpenThaiEval76.0574.1777.59
AVERAGE71.5863.3165.42

Technical report

OpenThaiGPT 1.6 and R1 Technical Report — arXivarxiv.org

If OpenThaiGPT has been beneficial for your work, kindly consider citing it as follows:

@misc{yuenyong2025openthaigpt16r1thaicentric,
title={OpenThaiGPT 1.6 and R1: Thai-Centric Open Source and Reasoning Large Language Models},
author={Sumeth Yuenyong and Thodsaporn Chay-intr and Kobkrit Viriyayudhakorn},
year={2025},
eprint={2504.01789},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2504.01789},
}

How to use

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "openthaigpt/openthaigpt-r1-32b-instruct"

model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)

prompt = "กรุงเทพมหานครคืออะไร"
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

generated_ids = model.generate(
**model_inputs,
max_new_tokens=8192,
temperature=0.6
)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]

response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]

vLLM

  1. Install vLLM (vllm-project/vllm)
  2. Run server
vllm serve openthaigpt/openthaigpt-r1-32b-instruct --tensor-parallel-size 2
  • Note, change --tensor-parallel-size 2 to the amount of available GPU cards.
  1. Run inference (CURL example)
curl -X POST 'http://127.0.0.1:8000/v1/chat/completions' \
-H 'Content-Type: application/json' \
-d '{
"model": "openthaigpt/openthaigpt-r1-32b-instruct",
"messages": [
{
"role": "user",
"content": "กรุงเทพมหานครคืออะไร"
}
],
"max_tokens": 4096,
"temperature": 0.6,
"top_p": 0.95,
"top_k": 40
}'

GPU memory requirements

Number of ParametersFP 16 bits8 bits (Quantized)4 bits (Quantized)
32b64 GB32 GB16 GB

License

  • This model is available for research and commercial use under the terms of the Qwen2.5 license agreement. Please refer to the LICENSE file for more information.

Support

Disclaimer: Provided responses are not guaranteed.