Git Analytics LLM & RAGLLM & RAG
lyogavin/airllm

AirLLM: การรันโมเดลภาษา LLM บน GPU 4GBAirLLM: Running LLM Models on a 4GB GPU

LLM & RAGLLM & RAG 35,442 ดาวStars 3,748 Jupyter Notebook Apache-2.0 อัปเดตล่าสุดLast push 6 ต.ค. 25696 Oct 2026
lyogavin/airllm

AirLLM ช่วยให้โมเดล LLM ขนาดใหญ่รันบน GPU 4GB โดยไม่ต้องใช้การลดขนาดหรือการบีบอัด

AirLLM allows large LLMs to run on a 4GB GPU without the need for size reduction or compression.

ไว้ทำอะไร

AirLLM ถูกออกแบบมาเพื่อลดการใช้งานหน่วยความจำของการ inference ในโมเดลภาษา LLM ขนาดใหญ่ ทำให้สามารถรันโมเดลที่ซับซ้อน เช่น Kimi K3 หรือ Qwen3.8-Flash-Next บน GPU ขนาดเล็กได้

ทำงานอย่างไร

AirLLM มีการติดตั้งโดยใช้แพ็คเกจ pip และการใช้งานผ่าน API ที่คล้ายกับ transformers ทั่วไป แต่เพิ่มความสามารถในการประมวลผลโมเดลขนาดใหญ่ผ่านการแบ่งเลเยอร์และการ streaming น้ำหนักโมเดลระหว่างการฝึกอบรมและการรัน inference. การกำหนดค่ามีการใช้ block-wise quantization เพื่อความเร็วในการ inference ที่สูงขึ้น.

โครงสร้างโค้ด

  • air_llm/airllm — โมดูลหลักสำหรับการจัดการโมเดล
  • air_llm/inference_example.py — ตัวอย่างการใช้งาน inference
  • anima_100k/ — ข้อมูลและการฝึกโมเดล
  • assets/ — ไฟล์ภาพและโลโก้
  • rlhf/ — การเรียนรู้ด้วยการเสริมกำลัง
  • training/ — สคริปต์ที่เกี่ยวข้องกับการฝึก

เริ่มใช้งาน

pip install airllm
from airllm import AutoModel

model = AutoModel.from_pretrained("Qwen/Qwen3-32B")

input_text = ['What is the capital of United States?']
input_tokens = model.tokenizer(input_text, return_tensors="pt", truncation=True)
generation_output = model.generate(input_tokens['input_ids'].cuda(), max_new_tokens=20)
output = model.tokenizer.decode(generation_output.sequences[0])
print(output)

เหมาะกับงานแบบไหน

  • รันโมเดล LLM ขนาดใหญ่บน GPU ขนาดเล็ก
  • การฝึกอบรมโมเดลภาษา AI ที่ซับซ้อนโดยใช้ทรัพยากรที่จำกัด

ข้อควรรู้

  • การใช้งานหนักอาจมีข้อจำกัดด้านประสิทธิภาพและความเร็วขึ้นอยู่กับความสามารถของฮาร์ดแวร์ที่ใช้
  • ใช้ License Apache-2.0 ทำให้เหมาะสมกับการใช้งานและการแจกจ่ายที่กว้าง แต่ต้องปฏิบัติตามเงื่อนไขของ License

What it is for

AirLLM is designed to reduce the inference memory usage of large LLMs, enabling complex models such as Kimi K3 or Qwen3.8-Flash-Next to run on smaller GPUs.

How it works

AirLLM is installed using a pip package and is utilized through an API similar to standard transformers, but adds capabilities for handling large models through layer partitioning and streaming model weights during training and inference. Configuration includes block-wise quantization for increased inference speed.

Code structure

  • air_llm/airllm — Main module for model management
  • air_llm/inference_example.py — Inference usage example
  • anima_100k/ — Data and training models
  • assets/ — Image files and logos
  • rlhf/ — Reinforcement learning resources
  • training/ — Training-related scripts

Getting started

pip install airllm
from airllm import AutoModel

model = AutoModel.from_pretrained("Qwen/Qwen3-32B")

input_text = ['What is the capital of United States?']
input_tokens = model.tokenizer(input_text, return_tensors="pt", truncation=True)
generation_output = model.generate(input_tokens['input_ids'].cuda(), max_new_tokens=20)
output = model.tokenizer.decode(generation_output.sequences[0])
print(output)

Good fit for

  • Running large LLM models on small GPUs
  • Training complex AI language models with limited resources

Things to know

  • Heavy usage may face performance and speed limitations depending on the hardware used
  • Uses Apache-2.0 license, which is suitable for widespread use and distribution but requires adherence to the license terms

บทวิเคราะห์นี้สร้างจาก README และโค้ดของ repo โดย AI ของ Oneable — ตรวจสอบ license และเอกสารต้นทางก่อนนำไปใช้งานจริงThis breakdown was generated from the repository's README and code by Oneable's AI — check the license and upstream docs before using it in production.

#chinese-llm#chinese-nlp#finetune#generative-ai#instruct-gpt#instruction-set#llama#llm#lora#open-models#open-source#open-source-models

repo อื่นในหมวดเดียวกันMore in this category