MobiLlama入门学习资料 - 面向边缘设备的小型语言模型

MobiLlama简介

MobiLlama是由阿联酋穆罕默德·本·扎耶德人工智能大学(MBZUAI)开发的一个小型语言模型(SLM),主要面向资源受限的边缘计算设备。与大型语言模型(LLM)不同,MobiLlama采用"小而美"的设计理念,在保持良好性能的同时,大幅降低了计算资源需求,非常适合在移动设备、可穿戴设备等边缘设备上运行。

MobiLlama的主要特点包括:

开源透明:完全开源的0.5B参数模型
高效轻量:针对边缘设备优化,资源需求低
性能出色:在多项基准测试中超越同等规模模型
灵活易用:支持多种规模版本,可根据需求选择

MobiLlama logo

模型下载

MobiLlama提供了多个版本的预训练模型供下载使用:

模型名称	下载链接
MobiLlama-05B	HuggingFace
MobiLlama-08B	HuggingFace
MobiLlama-1B	HuggingFace
MobiLlama-05B-Chat	HuggingFace
MobiLlama-1B-Chat	HuggingFace

快速使用

以下是使用MobiLlama进行文本生成的示例代码:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("MBZUAI/MobiLlama-05B", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("MBZUAI/MobiLlama-05B", trust_remote_code=True)

model.to('cuda')
text = "I was walking towards the river when "
input_ids = tokenizer(text, return_tensors="pt").to('cuda').input_ids
outputs = model.generate(input_ids, max_length=1000, repetition_penalty=1.2, pad_token_id=tokenizer.eos_token_id)
print(tokenizer.batch_decode(outputs[:, input_ids.shape[1]:-1])[0].strip())