MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co api & LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 github AI Model

Introduction of MoE-LLaVA-Phi2-2.7B-4e-384

Model Details of MoE-LLaVA-Phi2-2.7B-4e-384

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

If you like our project, please give us a star ⭐ on GitHub for latest update.

📰 News

[2024.01.30] The paper is released.
[2024.01.27] 🤗 Hugging Face demo and all codes & datasets are available now! Welcome to watch 👀 this repository for the latest updates.

😮 Highlights

MoE-LLaVA shows excellent performance in multi-modal learning.

🔥 High performance, but with fewer parameters

with just 3B sparsely activated parameters , MoE-LLaVA demonstrates performance comparable to the LLaVA-1.5-7B on various visual understanding datasets and even surpasses the LLaVA-1.5-13B in object hallucination benchmarks.

🚀 Simple baseline, learning multi-modal interactions with sparse pathways.

With the addition of a simple MoE tuning stage , we can complete the training of MoE-LLaVA on 8 V100 GPUs within 2 days.

🤗 Demo

Gradio Web UI

Highly recommend trying out our web demo by the following command, which incorporates all features currently supported by MoE-LLaVA. We also provide online demo in Huggingface Spaces.

# use phi2
deepspeed --include localhost:0 moellava/serve/gradio_web_server.py --model-path "LanguageBind/MoE-LLaVA-Phi2-2.7B-4e" 
# use qwen
deepspeed --include localhost:0 moellava/serve/gradio_web_server.py --model-path "LanguageBind/MoE-LLaVA-Qwen-1.8B-4e" 
# use stablelm
deepspeed --include localhost:0 moellava/serve/gradio_web_server.py --model-path "LanguageBind/MoE-LLaVA-StableLM-1.6B-4e"

CLI Inference

# use phi2
deepspeed --include localhost:0 moellava/serve/cli.py --model-path "LanguageBind/MoE-LLaVA-Phi2-2.7B-4e"  --image-file "image.jpg"
# use qwen
deepspeed --include localhost:0 moellava/serve/cli.py --model-path "LanguageBind/MoE-LLaVA-Qwen-1.8B-4e"  --image-file "image.jpg"
# use stablelm
deepspeed --include localhost:0 moellava/serve/cli.py --model-path "LanguageBind/MoE-LLaVA-StableLM-1.6B-4e"  --image-file "image.jpg"

🐳 Model Zoo

Model	LLM	Checkpoint	Avg	VQAv2	GQA	VizWiz	SQA	T-VQA	POPE	MM-Bench	LLaVA-Bench-Wild	MM-Vet
MoE-LLaVA-1.6B×4-Top2	1.6B	LanguageBind/MoE-LLaVA-StableLM-1.6B-4e	60.0	76.0	60.4	37.2	62.6	47.8	84.3	59.4	85.9	26.1
MoE-LLaVA-1.8B×4-Top2	1.8B	LanguageBind/MoE-LLaVA-Qwen-1.8B-4e	60.2	76.2	61.5	32.6	63.1	48.0	87.0	59.6	88.7	25.3
MoE-LLaVA-2.7B×4-Top2	2.7B	LanguageBind/MoE-LLaVA-Phi2-2.7B-4e	63.9	77.1	61.1	43.4	68.7	50.2	85.0	65.5	93.2	31.1

⚙️ Requirements and Installation

Python >= 3.10
Pytorch == 2.0.1
CUDA Version >= 11.7
Transformers == 4.36.2
Tokenizers==0.15.1
Install required packages:

git clone https://github.com/PKU-YuanGroup/MoE-LLaVA
cd MoE-LLaVA
conda create -n moellava python=3.10 -y
conda activate moellava
pip install --upgrade pip  # enable PEP 660 support
pip install -e .
pip install -e ".[train]"
pip install flash-attn --no-build-isolation

# Below are optional. For Qwen model.
git clone https://github.com/Dao-AILab/flash-attention
cd flash-attention && pip install .
# Below are optional. Installing them might be slow.
# pip install csrc/layer_norm
# If the version of flash-attn is higher than 2.1.1, the following is not needed.
# pip install csrc/rotary

🗝️ Training & Validating

The training & validating instruction is in TRAIN.md & EVAL.md .

💡 Customizing your MoE-LLaVA

The instruction is in CUSTOM.md .

😍 Visualization

The instruction is in VISUALIZATION.md .

🤖 API

We open source all codes. If you want to load the model (e.g. LanguageBind/MoE-LLaVA ) on local, you can use the following code snippets.

Using the following command to run the code.

deepspeed predict.py

import torch
from moellava.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKEN
from moellava.conversation import conv_templates, SeparatorStyle
from moellava.model.builder import load_pretrained_model
from moellava.utils import disable_torch_init
from moellava.mm_utils import tokenizer_image_token, get_model_name_from_path, KeywordsStoppingCriteria

def main():
    disable_torch_init()
    image = 'moellava/serve/examples/extreme_ironing.jpg'
    inp = 'What is unusual about this image?'
    model_path = 'LanguageBind/MoE-LLaVA-Phi2-2.7B-4e'  # LanguageBind/MoE-LLaVA-Qwen-1.8B-4e or LanguageBind/MoE-LLaVA-StableLM-1.6B-4e
    device = 'cuda'
    load_4bit, load_8bit = False, False  # FIXME: Deepspeed support 4bit or 8bit?
    model_name = get_model_name_from_path(model_path)
    tokenizer, model, processor, context_len = load_pretrained_model(model_path, None, model_name, load_8bit, load_4bit, device=device)
    image_processor = processor['image']
    conv_mode = "phi"  # qwen or stablelm
    conv = conv_templates[conv_mode].copy()
    roles = conv.roles
    image_tensor = image_processor.preprocess(image, return_tensors='pt')['pixel_values'].to(model.device, dtype=torch.float16)

    print(f"{roles[1]}: {inp}")
    inp = DEFAULT_IMAGE_TOKEN + '\n' + inp
    conv.append_message(conv.roles[0], inp)
    conv.append_message(conv.roles[1], None)
    prompt = conv.get_prompt()
    input_ids = tokenizer_image_token(prompt, tokenizer, IMAGE_TOKEN_INDEX, return_tensors='pt').unsqueeze(0).cuda()
    stop_str = conv.sep if conv.sep_style != SeparatorStyle.TWO else conv.sep2
    keywords = [stop_str]
    stopping_criteria = KeywordsStoppingCriteria(keywords, tokenizer, input_ids)

    with torch.inference_mode():
        output_ids = model.generate(
            input_ids,
            images=image_tensor,
            do_sample=True,
            temperature=0.2,
            max_new_tokens=1024,
            use_cache=True,
            stopping_criteria=[stopping_criteria])

    outputs = tokenizer.decode(output_ids[0, input_ids.shape[1]:], skip_special_tokens=True).strip()
    print(outputs)

if __name__ == '__main__':
    main()

🙌 Related Projects

Video-LLaVA This framework empowers the model to efficiently utilize the united visual tokens.
LanguageBind An open source five modalities language-based retrieval framework.

👍 Acknowledgement

LLaVA The codebase we built upon and it is an efficient large language and vision assistant.

🔒 License

The majority of this project is released under the Apache 2.0 license as found in the LICENSE file.
The service is a research preview intended for non-commercial use only, subject to the model License of LLaMA, Terms of Use of the data generated by OpenAI, and Privacy Practices of ShareGPT. Please contact us if you find any potential violation.

✏️ Citation

If you find our paper and code useful in your research, please consider giving a star :star: and citation :pencil:.

@misc{lin2024moellava,
      title={MoE-LLaVA: Mixture of Experts for Large Vision-Language Models}, 
      author={Bin Lin and Zhenyu Tang and Yang Ye and Jiaxi Cui and Bin Zhu and Peng Jin and Junwu Zhang and Munan Ning and Li Yuan},
      year={2024},
      eprint={2401.15947},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

@article{lin2023video,
  title={Video-LLaVA: Learning United Visual Representation by Alignment Before Projection},
  author={Lin, Bin and Zhu, Bin and Ye, Yang and Ning, Munan and Jin, Peng and Yuan, Li},
  journal={arXiv preprint arXiv:2311.10122},
  year={2023}
}

✨ Star History

🤝 Contributors

Runs of LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 on huggingface.co

1.0K

Total runs

24-hour runs

3-day runs

7-day runs

594

30-day runs

More Information About MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co Model

More MoE-LLaVA-Phi2-2.7B-4e-384 license Visit here:

https://choosealicense.com/licenses/apache-2.0

MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co

MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co is an AI model on huggingface.co that provides MoE-LLaVA-Phi2-2.7B-4e-384's model effect (), which can be used instantly with this LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 model. huggingface.co supports a free trial of the MoE-LLaVA-Phi2-2.7B-4e-384 model, and also provides paid use of the MoE-LLaVA-Phi2-2.7B-4e-384. Support call MoE-LLaVA-Phi2-2.7B-4e-384 model through api, including Node.js, Python, http.

MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co Url

https://huggingface.co/LanguageBind/MoE-LLaVA-Phi2-2.7B-4e-384

LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 online free

MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co is an online trial and call api platform, which integrates MoE-LLaVA-Phi2-2.7B-4e-384's modeling effects, including api services, and provides a free online trial of MoE-LLaVA-Phi2-2.7B-4e-384, you can try MoE-LLaVA-Phi2-2.7B-4e-384 online for free by clicking the link below.

LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 online free url in huggingface.co:

https://huggingface.co/LanguageBind/MoE-LLaVA-Phi2-2.7B-4e-384

MoE-LLaVA-Phi2-2.7B-4e-384 install

MoE-LLaVA-Phi2-2.7B-4e-384 is an open source model from GitHub that offers a free installation service, and any user can find MoE-LLaVA-Phi2-2.7B-4e-384 on GitHub to install. At the same time, huggingface.co provides the effect of MoE-LLaVA-Phi2-2.7B-4e-384 install, users can directly use MoE-LLaVA-Phi2-2.7B-4e-384 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

MoE-LLaVA-Phi2-2.7B-4e-384 install url in huggingface.co:

https://huggingface.co/LanguageBind/MoE-LLaVA-Phi2-2.7B-4e-384

huggingface.co

LanguageBind/LanguageBind_Image

Total runs: 196.0K

Run Growth: 162.6K

Growth Rate: 83.00%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Video_merge

Total runs: 192.1K

Run Growth: 150.6K

Growth Rate: 82.24%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Video_FT

Total runs: 27.9K

Run Growth: -78.0K

Growth Rate: -272.90%

Updated: February 01 2024

huggingface.co

LanguageBind/Video-LLaVA-7B

Total runs: 18.7K

Run Growth: 6.5K

Growth Rate: 35.13%

Updated: April 09 2024

huggingface.co

LanguageBind/Video-LLaVA-7B-hf

Total runs: 15.9K

Run Growth: -4.4K

Growth Rate: -27.41%

Updated: May 16 2024

huggingface.co

LanguageBind/LanguageBind_Audio_FT

Total runs: 4.8K

Run Growth: 329

Growth Rate: 7.03%

Updated: February 01 2024

huggingface.co

LanguageBind/MoE-LLaVA-StableLM-1.6B-4e

Total runs: 3.4K

Run Growth: 1.7K

Growth Rate: 49.88%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Video_V1.5_FT

Total runs: 1.8K

Run Growth: -709

Growth Rate: -41.68%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Audio

Total runs: 470

Run Growth: 190

Growth Rate: 40.43%

Updated: February 01 2024

huggingface.co

LanguageBind/MoE-LLaVA-Phi2-2.7B-4e

Total runs: 377

Run Growth: 136

Growth Rate: 36.07%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Video

Total runs: 345

Run Growth: -113

Growth Rate: -32.56%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Video_Huge_V1.5_FT

Total runs: 269

Run Growth: -314

Growth Rate: -101.95%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Thermal

Total runs: 195

Run Growth: -73

Growth Rate: -37.44%

Updated: February 01 2024

huggingface.co

LanguageBind/LanguageBind_Depth

Total runs: 164

Run Growth: -133

Growth Rate: -82.10%

Updated: February 01 2024

huggingface.co

LanguageBind/MoE-LLaVA-Qwen-1.8B-4e

Total runs: 107

Run Growth: -166

Growth Rate: -155.14%

Updated: February 01 2024

huggingface.co

LanguageBind/MoE-LLaVA-StableLM-1.6B-4e-384

Total runs: 51

Run Growth: -36

Growth Rate: -70.59%

Updated: February 03 2024

huggingface.co

LanguageBind/Video-LLaVA-Pretrain-7B

Total runs: 27

Run Growth: -2

Growth Rate: -7.69%

Updated: February 01 2024

huggingface.co

LanguageBind/MoE-LLaVA-Phi2-Stage2

Total runs: 25

Run Growth: 4

Growth Rate: 16.00%

Updated: March 16 2024

huggingface.co

LanguageBind/MoE-LLaVA-Qwen-Pretrain

Total runs: 20

Run Growth: 9

Growth Rate: 45.00%

Updated: February 01 2024

huggingface.co

LanguageBind/MoE-LLaVA-Qwen-Stage2

Total runs: 18

Run Growth: -7

Growth Rate: -63.64%

Updated: March 16 2024

huggingface.co

LanguageBind/MoE-LLaVA-StableLM-Stage2-384

Total runs: 7

Run Growth: -8

Growth Rate: -114.29%

Updated: March 16 2024

huggingface.co

LanguageBind/MoE-LLaVA-Phi2-Stage2-384

Total runs: 6

Run Growth: -22

Growth Rate: -366.67%

Updated: March 16 2024

huggingface.co

LanguageBind/MoE-LLaVA-OpenChat-7B-4e

Total runs: 4

Run Growth: -9

Growth Rate: -225.00%

Updated: February 02 2024

huggingface.co

LanguageBind/MoE-LLaVA-StableLM-384-Pretrain

Total runs: 3

Run Growth: -6

Growth Rate: -200.00%

Updated: March 16 2024

huggingface.co

LanguageBind/MoE-LLaVA-StableLM-Stage2

Total runs: 3

Run Growth: -14

Growth Rate: -466.67%

Updated: March 16 2024

huggingface.co

LanguageBind/Open-Sora-Plan-v1.3.0

Total runs: 3

Run Growth: 4

Growth Rate: 22.22%

Updated: December 05 2024

huggingface.co

LanguageBind/Open-Sora-Plan-v1.1.0

Total runs: 2

Run Growth: 1

Growth Rate: 50.00%

Updated: May 27 2024

huggingface.co

LanguageBind/Open-Sora-Plan-v1.0.0

Total runs: 2

Run Growth: 0

Growth Rate: 0.00%

Updated: April 07 2024

huggingface.co

LanguageBind/MoE-LLaVA-Phi2-Pretrain

Total runs: 2

Run Growth: -11

Growth Rate: -550.00%

Updated: February 01 2024

huggingface.co

LanguageBind/MoE-LLaVA-Phi2-384-Pretrain

Total runs: 2

Run Growth: -14

Growth Rate: -700.00%

Updated: February 01 2024

huggingface.co

LanguageBind/Open-Sora-Plan-v1.2.0

Total runs: 2

Run Growth: 0

Growth Rate: 0.00%

Updated: September 07 2024

huggingface.co

LanguageBind/MoE-LLaVA-StableLM-Pretrain

Total runs: 1

Run Growth: -10

Growth Rate: -1000.00%

Updated: February 01 2024

huggingface.co

LanguageBind/Video-LLaVA-V1.5

Total runs: 0

Run Growth: 0

Growth Rate: 0.00%

Updated: November 26 2023

huggingface.co

LanguageBind/offline_feature

Total runs: 0

Run Growth: 0

Growth Rate: 0.00%

Updated: January 29 2025

huggingface.co

LanguageBind/LanguageBind_Audio_V1.5

Total runs: 0

Run Growth: 0

Growth Rate: 0.00%

Updated: February 01 2024

LanguageBind / MoE-LLaVA-Phi2-2.7B-4e-384

Introduction of MoE-LLaVA-Phi2-2.7B-4e-384

Model Details of MoE-LLaVA-Phi2-2.7B-4e-384

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

If you like our project, please give us a star ⭐ on GitHub for latest update.

📰 News

😮 Highlights

🔥 High performance, but with fewer parameters

🚀 Simple baseline, learning multi-modal interactions with sparse pathways.

🤗 Demo

Gradio Web UI

CLI Inference

🐳 Model Zoo

⚙️ Requirements and Installation

🗝️ Training & Validating

💡 Customizing your MoE-LLaVA

😍 Visualization

🤖 API

🙌 Related Projects

👍 Acknowledgement

🔒 License

✏️ Citation

✨ Star History

🤝 Contributors

Runs of LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 on huggingface.co

More Information About MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co Model

More MoE-LLaVA-Phi2-2.7B-4e-384 license Visit here:

MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co

MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co Url

LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 online free

LanguageBind MoE-LLaVA-Phi2-2.7B-4e-384 online free url in huggingface.co:

MoE-LLaVA-Phi2-2.7B-4e-384 install

MoE-LLaVA-Phi2-2.7B-4e-384 install url in huggingface.co:

Url of MoE-LLaVA-Phi2-2.7B-4e-384

MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co Url

Provider of MoE-LLaVA-Phi2-2.7B-4e-384 huggingface.co

Other API from LanguageBind