Instructions to use BK-Lee/MoAI-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BK-Lee/MoAI-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="BK-Lee/MoAI-7B")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("BK-Lee/MoAI-7B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use BK-Lee/MoAI-7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BK-Lee/MoAI-7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BK-Lee/MoAI-7B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/BK-Lee/MoAI-7B
- SGLang
How to use BK-Lee/MoAI-7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BK-Lee/MoAI-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BK-Lee/MoAI-7B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BK-Lee/MoAI-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BK-Lee/MoAI-7B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use BK-Lee/MoAI-7B with Docker Model Runner:
docker model run hf.co/BK-Lee/MoAI-7B
|
Download README.md from BK-Lee/MoAI-7B: direct link, hf CLI and curl.
- Browser
- Download file 2.23 kB
-
https://huggingface.co/BK-Lee/MoAI-7B/resolve/main/README.md
- Command line
-
hf download hf://BK-Lee/MoAI-7B/README.md
-
curl -L -o README.md https://huggingface.co/BK-Lee/MoAI-7B/resolve/main/README.md
2.23 kB
| license: mit | |
| pipeline_tag: image-text-to-text | |
| ## MoAI model | |
| This repository contains the weights of the model presented in [MoAI: Mixture of All Intelligence for Large Language and Vision Models](https://huggingface.co/papers/2403.07508). | |
| ### Simple running code is based on [MoAI-Github](https://github.com/ByungKwanLee/MoAI). | |
| You need only the following seven steps. | |
| ### [0] Download Github Code of MoAI, install the required libraries, set the necessary environment variable (README.md explains in detail! Don't Worry!). | |
| ```bash | |
| git clone https://github.com/ByungKwanLee/MoAI | |
| bash install | |
| ``` | |
| ### [1] Loading Image | |
| ```python | |
| from PIL import Image | |
| from torchvision.transforms import Resize | |
| from torchvision.transforms.functional import pil_to_tensor | |
| image_path = "figures/moai_mystery.png" | |
| image = Resize(size=(490, 490), antialias=False)(pil_to_tensor(Image.open(image_path))) | |
| ``` | |
| ### [2] Instruction Prompt | |
| ```python | |
| prompt = "Describe this image in detail." | |
| ``` | |
| ### [3] Loading MoAI | |
| ```python | |
| from moai.load_moai import prepare_moai | |
| moai_model, moai_processor, seg_model, seg_processor, od_model, od_processor, sgg_model, ocr_model \ | |
| = prepare_moai(moai_path='BK-Lee/MoAI-7B', bits=4, grad_ckpt=False, lora=False, dtype='fp16') | |
| ``` | |
| ### [4] Pre-processing for MoAI | |
| ```python | |
| moai_inputs = moai_model.demo_process(image=image, | |
| prompt=prompt, | |
| processor=moai_processor, | |
| seg_model=seg_model, | |
| seg_processor=seg_processor, | |
| od_model=od_model, | |
| od_processor=od_processor, | |
| sgg_model=sgg_model, | |
| ocr_model=ocr_model, | |
| device='cuda:0') | |
| ``` | |
| ### [5] Generate | |
| ```python | |
| import torch | |
| with torch.inference_mode(): | |
| generate_ids = moai_model.generate(**moai_inputs, do_sample=True, temperature=0.9, top_p=0.95, max_new_tokens=256, use_cache=True) | |
| ``` | |
| ### [6] Decoding | |
| ```python | |
| answer = moai_processor.batch_decode(generate_ids, skip_special_tokens=True)[0].split('[U')[0] | |
| print(answer) | |
| ``` |