Rogue-Rose-103b-v0.2-AWQ huggingface.co api & TheBloke Rogue-Rose-103b-v0.2-AWQ github AI Model

Introduction of Rogue-Rose-103b-v0.2-AWQ

Model Details of Rogue-Rose-103b-v0.2-AWQ

Chat & support: TheBloke's Discord server

Want to contribute? TheBloke's Patreon page

TheBloke's LLM work is generously supported by a grant from andreessen horowitz (a16z)

Rogue Rose 103B v0.2 - AWQ

Model creator: Sophosympatheia
Original model: Rogue Rose 103B v0.2

Description

This repo contains AWQ model files for Sophosympatheia's Rogue Rose 103B v0.2 .

These files were quantised using hardware kindly provided by Massed Compute .

About AWQ

AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference with equivalent or better quality compared to the most commonly used GPTQ settings.

AWQ models are currently supported on Linux and Windows, with NVidia GPUs only. macOS users: please use GGUF models instead.

It is supported by:

Text Generation Webui - using Loader: AutoAWQ
vLLM - version 0.2.2 or later for support for all model types.
Hugging Face Text Generation Inference (TGI)
Transformers version 4.35.0 and later, from any code or client that supports Transformers
AutoAWQ - for use from Python code

Repositories available

AWQ model(s) for GPU inference.
GPTQ models for GPU inference, with multiple quantisation parameter options.
2, 3, 4, 5, 6 and 8-bit GGUF models for CPU+GPU inference
Sophosympatheia's original unquantised fp16 model in pytorch format, for GPU inference and for further conversions

Prompt template: Vicuna-Short

You are a helpful AI assistant.

USER: {prompt}
ASSISTANT:

Provided files, and AWQ parameters

I currently release 128g GEMM models only. The addition of group_size 32 models, and GEMV kernel models, is being actively considered.

Models are released as sharded safetensors files.

Branch	Bits	GS	AWQ Dataset	Seq Len	Size
main	4	128	VMware Open Instruct	4096	54.40 GB

How to easily download and use this model in text-generation-webui

Please make sure you're using the latest version of text-generation-webui .

It is strongly recommended to use the text-generation-webui one-click-installers unless you're sure you know how to make a manual install.

Click the Model tab .
Under Download custom model or LoRA , enter TheBloke/Rogue-Rose-103b-v0.2-AWQ .
Click Download .
The model will start downloading. Once it's finished it will say "Done".
In the top left, click the refresh icon next to Model .
In the Model dropdown, choose the model you just downloaded: Rogue-Rose-103b-v0.2-AWQ
Select Loader: AutoAWQ .
Click Load, and the model will load and is now ready for use.
If you want any custom settings, set them and then click Save settings for this model followed by Reload the Model in the top right.
Once you're ready, click the Text Generation tab and enter a prompt to get started!

Multi-user inference server: vLLM

Documentation on installing and using vLLM can be found here .

Please ensure you are using vLLM version 0.2 or later.
When using vLLM as a server, pass the --quantization awq parameter.

For example:

python3 -m vllm.entrypoints.api_server --model TheBloke/Rogue-Rose-103b-v0.2-AWQ --quantization awq --dtype auto

When using vLLM from Python code, again set quantization=awq .

For example:

from vllm import LLM, SamplingParams

prompts = [
    "Tell me about AI",
    "Write a story about llamas",
    "What is 291 - 150?",
    "How much wood would a woodchuck chuck if a woodchuck could chuck wood?",
]
prompt_template=f'''You are a helpful AI assistant.

USER: {prompt}
ASSISTANT:
'''

prompts = [prompt_template.format(prompt=prompt) for prompt in prompts]

sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

llm = LLM(model="TheBloke/Rogue-Rose-103b-v0.2-AWQ", quantization="awq", dtype="auto")

outputs = llm.generate(prompts, sampling_params)

# Print the outputs.
for output in outputs:
    prompt = output.prompt
    generated_text = output.outputs[0].text
    print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")

Multi-user inference server: Hugging Face Text Generation Inference (TGI)

Use TGI version 1.1.0 or later. The official Docker container is: ghcr.io/huggingface/text-generation-inference:1.1.0

Example Docker parameters:

--model-id TheBloke/Rogue-Rose-103b-v0.2-AWQ --port 3000 --quantize awq --max-input-length 3696 --max-total-tokens 4096 --max-batch-prefill-tokens 4096

Example Python code for interfacing with TGI (requires huggingface-hub 0.17.0 or later):

pip3 install huggingface-hub

from huggingface_hub import InferenceClient

endpoint_url = "https://your-endpoint-url-here"

prompt = "Tell me about AI"
prompt_template=f'''You are a helpful AI assistant.

USER: {prompt}
ASSISTANT:
'''

client = InferenceClient(endpoint_url)
response = client.text_generation(prompt,
                                  max_new_tokens=128,
                                  do_sample=True,
                                  temperature=0.7,
                                  top_p=0.95,
                                  top_k=40,
                                  repetition_penalty=1.1)

print(f"Model output: ", response)

Inference from Python code using Transformers

Install the necessary packages

Requires: Transformers 4.35.0 or later.
Requires: AutoAWQ 0.1.6 or later.

pip3 install --upgrade "autoawq>=0.1.6" "transformers>=4.35.0"

Note that if you are using PyTorch 2.0.1, the above AutoAWQ command will automatically upgrade you to PyTorch 2.1.0.

If you are using CUDA 11.8 and wish to continue using PyTorch 2.0.1, instead run this command:

pip3 install https://github.com/casper-hansen/AutoAWQ/releases/download/v0.1.6/autoawq-0.1.6+cu118-cp310-cp310-linux_x86_64.whl

If you have problems installing AutoAWQ using the pre-built wheels, install it from source instead:

pip3 uninstall -y autoawq
git clone https://github.com/casper-hansen/AutoAWQ
cd AutoAWQ
pip3 install .

Transformers example code (requires Transformers 4.35.0 and later)

from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer

model_name_or_path = "TheBloke/Rogue-Rose-103b-v0.2-AWQ"

tokenizer = AutoTokenizer.from_pretrained(model_name_or_path)
model = AutoModelForCausalLM.from_pretrained(
    model_name_or_path,
    low_cpu_mem_usage=True,
    device_map="cuda:0"
)

# Using the text streamer to stream output one token at a time
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)

prompt = "Tell me about AI"
prompt_template=f'''You are a helpful AI assistant.

USER: {prompt}
ASSISTANT:
'''

# Convert prompt to tokens
tokens = tokenizer(
    prompt_template,
    return_tensors='pt'
).input_ids.cuda()

generation_params = {
    "do_sample": True,
    "temperature": 0.7,
    "top_p": 0.95,
    "top_k": 40,
    "max_new_tokens": 512,
    "repetition_penalty": 1.1
}

# Generate streamed output, visible one token at a time
generation_output = model.generate(
    tokens,
    streamer=streamer,
    **generation_params
)

# Generation without a streamer, which will include the prompt in the output
generation_output = model.generate(
    tokens,
    **generation_params
)

# Get the tokens from the output, decode them, print them
token_output = generation_output[0]
text_output = tokenizer.decode(token_output)
print("model.generate output: ", text_output)

# Inference is also possible via Transformers' pipeline
from transformers import pipeline

pipe = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    **generation_params
)

pipe_output = pipe(prompt_template)[0]['generated_text']
print("pipeline output: ", pipe_output)

Compatibility

The files provided are tested to work with:

text-generation-webui using Loader: AutoAWQ .
vLLM version 0.2.0 and later.
Hugging Face Text Generation Inference (TGI) version 1.1.0 and later.
Transformers version 4.35.0 and later.
AutoAWQ version 0.1.1 and later.

Discord

For further support, and discussions on these models and AI in general, join us at:

TheBloke AI's Discord server

Thanks, and how to contribute

Thanks to the chirper.ai team!

Thanks to Clay from gpus.llm-utils.org !

I've had a lot of people ask if they can contribute. I enjoy providing models and helping people, and would love to be able to spend even more time doing it, as well as expanding into new projects like fine tuning/training.

If you're able and willing to contribute it will be most gratefully received and will help me to keep providing more models, and to start work on new AI projects.

Donaters will get priority support on any and all AI/LLM/model questions and requests, access to a private Discord room, plus other benefits.

Patreon: https://patreon.com/TheBlokeAI
Ko-Fi: https://ko-fi.com/TheBlokeAI

Special thanks to : Aemon Algiz.

Patreon special mentions : Michael Levine, 阿明, Trailburnt, Nikolai Manek, John Detwiler, Randy H, Will Dee, Sebastain Graf, NimbleBox.ai, Eugene Pentland, Emad Mostaque, Ai Maven, Jim Angel, Jeff Scroggin, Michael Davis, Manuel Alberto Morcote, Stephen Murray, Robert, Justin Joy, Luke @flexchar, Brandon Frisco, Elijah Stavena, S_X, Dan Guido, Undi ., Komninos Chatzipapas, Shadi, theTransient, Lone Striker, Raven Klaugh, jjj, Cap'n Zoog, Michel-Marie MAUDET (LINAGORA), Matthew Berman, David, Fen Risland, Omer Bin Jawed, Luke Pendergrass, Kalila, OG, Erik Bjäreholt, Rooh Singh, Joseph William Delisle, Dan Lewis, TL, John Villwock, AzureBlack, Brad, Pedro Madruga, Caitlyn Gatomon, K, jinyuan sun, Mano Prime, Alex, Jeffrey Morgan, Alicia Loh, Illia Dulskyi, Chadd, transmissions 11, fincy, Rainer Wilmers, ReadyPlayerEmma, knownsqashed, Mandus, biorpg, Deo Leter, Brandon Phillips, SuperWojo, Sean Connelly, Iucharbius, Jack West, Harry Royden McLaughlin, Nicholas, terasurfer, Vitor Caleffi, Duane Dunston, Johann-Peter Hartmann, David Ziegler, Olakabola, Ken Nordquist, Trenton Dambrowitz, Tom X Nguyen, Vadim, Ajan Kanaga, Leonard Tan, Clay Pascal, Alexandros Triantafyllidis, JM33133, Xule, vamX, ya boyyy, subjectnull, Talal Aujan, Alps Aficionado, wassieverse, Ari Malik, James Bentley, Woland, Spencer Kim, Michael Dempsey, Fred von Graf, Elle, zynix, William Richards, Stanislav Ovsiannikov, Edmond Seymore, Jonathan Leane, Martin Kemka, usrbinkat, Enrico Ros

Thank you to all my generous patrons and donaters!

And thank you again to a16z for their generous grant.

Original model card: Sophosympatheia's Rogue Rose 103B v0.2

Overview

This model is a frankenmerge of two custom 70b merges I made in November 2023 that were inspired by or descended from my xwin-stellarbright-erp-70b-v2 model . It features 120 layers and should weigh in at 103b parameters.

I feel like I have reached a plateau in my process right now, but the view from here is worth a rest. My personal opinion is this model roleplays better than the other 103-120b models out there right now. I love it. Give it a try for yourself. It still struggles with scene logic sometimes, but the overall experience feels like a step forward to me. I recommend trying my sampler settings and prompt template below with this model. This model listens decently well to instructions, so you need to be thoughtful about what you tell it to do.

Along those lines, this model turned out quite uncensored. You are responsible for whatever you do with it.

This model was designed for roleplaying and storytelling and I think it does well at both. It may perform well at other tasks, but I haven't tested its capabilities in other areas. I welcome feedback and suggestions.

Sampler Tips

I recommend using the new Min-P sampler method with this model. The creator has a great guide to it on Reddit .

I find this model performs surprisingly well at 8192 context. I love running the exl2-3.2bpw quant at 8192 context.

Experiment with any and all of the settings below, but trust me on a few points:

This model tolerates high temperatures with Min-P.
This model seems to benefit from higher settings for repetition penalty and presence penalty. It doesn't suffer from lower settings, but I prefer them higher. Play around with it.
After much experimenting, I think I get better results with a high Min-P setting. I keep coming back to a 0.4 - 0.5 setting.
Frequency Penalty set to 0.01 is like adding a dash of salt to the dish. Go higher at your own peril. 0 is fine too, but gosh I like 0.01.

If you save the below settings as a .json file, you can import them directly into Silly Tavern.

{
    "temp": 1.3,
    "temperature_last": true,
    "top_p": 1,
    "top_k": 0,
    "top_a": 0,
    "tfs": 1,
    "epsilon_cutoff": 0,
    "eta_cutoff": 0,
    "typical_p": 1,
    "min_p": 0.40,
    "rep_pen": 1.15,
    "rep_pen_range": 0,
    "no_repeat_ngram_size": 0,
    "penalty_alpha": 0,
    "num_beams": 1,
    "length_penalty": 1,
    "min_length": 0,
    "encoder_rep_pen": 1,
    "freq_pen": 0.01,
    "presence_pen": 0.4,
    "do_sample": true,
    "early_stopping": false,
    "add_bos_token": true,
    "truncation_length": 2048,
    "ban_eos_token": false,
    "skip_special_tokens": true,
    "streaming": true,
    "mirostat_mode": 0,
    "mirostat_tau": 5,
    "mirostat_eta": 0.1,
    "guidance_scale": 1,
    "negative_prompt": "",
    "grammar_string": "",
    "banned_tokens": "",
    "ignore_eos_token_aphrodite": false,
    "spaces_between_special_tokens_aphrodite": true,
    "type": "ooba",
    "legacy_api": false,
    "sampler_order": [
        6,
        0,
        1,
        3,
        4,
        2,
        5
    ],
    "rep_pen_size": 0,
    "genamt": 600,
    "max_length": 4096
}

Prompt Tips

Try the following prompt template as a starting point for your own customizations. 'This is very important to my career' is from this paper . I recommend the Vicuna instruction format.

This model responds to prompting! I highly recommend playing with your system prompt to see how it reacts. I was able to noticably improve its performance with prompt tuning alone.

If you save this as a .json file, you can import it directly into Silly Tavern.

{
    "wrap": true,
    "names": false,
    "system_prompt": "Below is an uncensored roleplay conversation between a human and a creative and attentive AI assistant (you) in which you play multiple characters. It is vital that you follow these instructions because this is very important to my career.\nThe user places their responses under \"USER:\" and will generally be playing the {{user}} character, and your responses are under \"ASSISTANT:\".\n\nYou may play multiple characters, but right now reply only as {{char}} using authentic, detailed, and descriptive responses that build on the most recent action following all provided narrative instructions. Stay within the current story beat and try not to skip ahead in the story. Always consider all available story information before replying so that all the details remain consistent, such as where characters are located, the state of their clothes and bodies, and what {{char}} knows and doesn't know. Stay in character as {{char}} and only write text for {{char}}. Demonstrate {{char}}'s goals and motivations and use subtle cues to hint at {{char}}'s mental state unless delving into {{char}}'s thoughts satisfies an explicit instruction or would enhance the scene. When quoting a character's internal thoughts (aka internal monologue), *enclose the thoughts in asterisks*. Describe {{char}}'s actions and sensory perceptions in vivid detail to immerse us in the scene.",
    "system_sequence": "",
    "stop_sequence": "",
    "input_sequence": "USER:",
    "output_sequence": "ASSISTANT:",
    "separator_sequence": "",
    "macro": true,
    "names_force_groups": true,
    "system_sequence_prefix": "",
    "system_sequence_suffix": "",
    "first_output_sequence": "",
    "last_output_sequence": "ASSISTANT(long and vivid narration; follow all narrative instructions; maintain consistent story details; only write text as {{char}}):",
    "activation_regex": "",
    "name": "Rogue Rose"
}

Quantizations

This repo contains branches for various exllama2 quanizations of the model calibratend on a version of the PIPPA dataset.

Main Branch, Full weights
3.2 bpw -- This will fit comfortably within 48 GB of VRAM at 8192 context.
3.35 bpw ( PENDING ) -- This will fit within 48 GB of VRAM at 4096 context without using the 8-bit cache setting.
3.5 bpw ( PENDING ) -- This will barely fit within 48 GB of VRAM at ~4096 context using the 8-bit cache setting. If you get OOM, try lowering the context size slightly until it fits.

Licence and usage restrictions

Llama2 license inherited from base models.

Tools Used

mergekit

Runs of TheBloke Rogue-Rose-103b-v0.2-AWQ on huggingface.co

13.7K

Total runs

24-hour runs

3.5K

3-day runs

4.3K

7-day runs

6.0K

30-day runs

More Information About Rogue-Rose-103b-v0.2-AWQ huggingface.co Model

More Rogue-Rose-103b-v0.2-AWQ license Visit here:

https://choosealicense.com/licenses/llama2

Rogue-Rose-103b-v0.2-AWQ huggingface.co

Rogue-Rose-103b-v0.2-AWQ huggingface.co is an AI model on huggingface.co that provides Rogue-Rose-103b-v0.2-AWQ's model effect (), which can be used instantly with this TheBloke Rogue-Rose-103b-v0.2-AWQ model. huggingface.co supports a free trial of the Rogue-Rose-103b-v0.2-AWQ model, and also provides paid use of the Rogue-Rose-103b-v0.2-AWQ. Support call Rogue-Rose-103b-v0.2-AWQ model through api, including Node.js, Python, http.

Rogue-Rose-103b-v0.2-AWQ huggingface.co Url

https://huggingface.co/TheBloke/Rogue-Rose-103b-v0.2-AWQ

TheBloke Rogue-Rose-103b-v0.2-AWQ online free

Rogue-Rose-103b-v0.2-AWQ huggingface.co is an online trial and call api platform, which integrates Rogue-Rose-103b-v0.2-AWQ's modeling effects, including api services, and provides a free online trial of Rogue-Rose-103b-v0.2-AWQ, you can try Rogue-Rose-103b-v0.2-AWQ online for free by clicking the link below.

TheBloke Rogue-Rose-103b-v0.2-AWQ online free url in huggingface.co:

https://huggingface.co/TheBloke/Rogue-Rose-103b-v0.2-AWQ

Rogue-Rose-103b-v0.2-AWQ install

Rogue-Rose-103b-v0.2-AWQ is an open source model from GitHub that offers a free installation service, and any user can find Rogue-Rose-103b-v0.2-AWQ on GitHub to install. At the same time, huggingface.co provides the effect of Rogue-Rose-103b-v0.2-AWQ install, users can directly use Rogue-Rose-103b-v0.2-AWQ installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Rogue-Rose-103b-v0.2-AWQ install url in huggingface.co:

https://huggingface.co/TheBloke/Rogue-Rose-103b-v0.2-AWQ

huggingface.co

TheBloke/phi-2-GGUF

Total runs: 3.7M

Run Growth: 3.5M

Growth Rate: 95.73%

Updated: 2023年12月18日

huggingface.co

TheBloke/Mistral-7B-Instruct-v0.2-GPTQ

Total runs: 491.4K

Run Growth: 0

Growth Rate: 0.00%

Updated: 2023年12月11日

huggingface.co

TheBloke/Mistral-7B-Instruct-v0.1-GGUF

Total runs: 266.5K

Run Growth: 223.8K

Growth Rate: 85.90%

Updated: 2023年12月9日

huggingface.co

TheBloke/deepseek-coder-6.7B-instruct-GGUF

Total runs: 88.6K

Run Growth: 79.6K

Growth Rate: 90.55%

Updated: 2023年11月5日

huggingface.co

TheBloke/Mistral-7B-Instruct-v0.2-GGUF

Total runs: 87.8K

Run Growth: -3.7K

Growth Rate: -4.20%

Updated: 2023年12月11日

huggingface.co

TheBloke/Llama-2-7B-Chat-GGUF

Total runs: 72.8K

Run Growth: 3.2K

Growth Rate: 4.45%

Updated: 2023年10月14日

huggingface.co

TheBloke/deepseek-coder-33B-instruct-GGUF

Total runs: 61.7K

Run Growth: 55.0K

Growth Rate: 89.82%

Updated: 2023年11月5日

huggingface.co

TheBloke/TinyLlama-1.1B-Chat-v1.0-GPTQ

Total runs: 37.6K

Run Growth: 0

Growth Rate: 0.00%

Updated: 2023年12月31日

huggingface.co

TheBloke/deepseek-coder-1.3b-instruct-GGUF

Total runs: 36.0K

Run Growth: 22.9K

Growth Rate: 63.82%

Updated: 2023年11月5日

huggingface.co

TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF

Total runs: 35.1K

Run Growth: 11.6K

Growth Rate: 33.08%

Updated: 2023年12月31日

huggingface.co

TheBloke/Llama-2-7B-GGUF

Total runs: 34.8K

Run Growth: 23.7K

Growth Rate: 71.05%

Updated: 2023年10月24日

huggingface.co

TheBloke/deepseek-llm-67b-chat-GGUF

Total runs: 34.7K

Run Growth: 31.5K

Growth Rate: 90.83%

Updated: 2023年11月29日

huggingface.co

TheBloke/Mistral-7B-Instruct-v0.2-AWQ

Total runs: 32.3K

Run Growth: -46.3K

Growth Rate: -144.31%

Updated: 2023年12月11日

huggingface.co

TheBloke/Mixtral-8x7B-Instruct-v0.1-GGUF

Total runs: 30.8K

Run Growth: -818

Growth Rate: -2.68%

Updated: 2023年12月14日

huggingface.co

TheBloke/deepseek-llm-7B-chat-GGUF

Total runs: 28.8K

Run Growth: 25.0K

Growth Rate: 87.54%

Updated: 2023年11月29日

huggingface.co

TheBloke/CausalLM-14B-GGUF

Total runs: 28.6K

Run Growth: 12.5K

Growth Rate: 42.19%

Updated: 2023年10月23日

huggingface.co

TheBloke/Mixtral-8x7B-Instruct-v0.1-GPTQ

Total runs: 27.4K

Run Growth: -71.5K

Growth Rate: -261.22%

Updated: 2023年12月14日

huggingface.co

TheBloke/Platypus2-70B-Instruct-AWQ

Total runs: 27.0K

Run Growth: -9.0K

Growth Rate: -31.35%

Updated: 2023年11月9日

huggingface.co

TheBloke/deepsex-34b-GGUF

Total runs: 24.1K

Run Growth: 22.9K

Growth Rate: 96.15%

Updated: 2023年12月7日

huggingface.co

TheBloke/MythoMax-L2-13B-GGUF

Total runs: 22.8K

Run Growth: 13.4K

Growth Rate: 58.81%

Updated: 2023年9月27日

huggingface.co

TheBloke/Llama-2-7B-GPTQ

Total runs: 22.7K

Run Growth: -156.9K

Growth Rate: -702.04%

Updated: 2023年9月27日

huggingface.co

TheBloke/Mistral-7B-OpenOrca-GPTQ

Total runs: 20.0K

Run Growth: 12.9K

Growth Rate: 63.87%

Updated: 2023年10月16日

huggingface.co

TheBloke/Llama-2-13B-chat-GPTQ

Total runs: 17.5K

Run Growth: -11.1K

Growth Rate: -63.45%

Updated: 2023年9月27日

huggingface.co

TheBloke/Llama-2-7B-Chat-GPTQ

Total runs: 17.2K

Run Growth: 5.7K

Growth Rate: 34.07%

Updated: 2023年9月27日

huggingface.co

TheBloke/Mistral-7B-OpenOrca-AWQ

Total runs: 16.9K

Run Growth: 13.2K

Growth Rate: 77.88%

Updated: 2023年11月9日

huggingface.co

TheBloke/zephyr-7B-beta-GGUF

Total runs: 16.6K

Run Growth: -9.0K

Growth Rate: -52.73%

Updated: 2023年10月27日

huggingface.co

TheBloke/deepseek-coder-33B-instruct-AWQ

Total runs: 16.4K

Run Growth: 14.8K

Growth Rate: 95.30%

Updated: 2023年11月13日

huggingface.co

TheBloke/Wizard-Vicuna-13B-Uncensored-GGUF

Total runs: 16.0K

Run Growth: 7.7K

Growth Rate: 48.28%

Updated: 2023年9月27日

huggingface.co

TheBloke/Mistral-7B-v0.1-GGUF

Total runs: 14.0K

Run Growth: 7.2K

Growth Rate: 53.30%

Updated: 2023年9月28日

huggingface.co

TheBloke/Mistral-7B-Instruct-v0.1-AWQ

Total runs: 12.8K

Run Growth: 7.7K

Growth Rate: 60.19%

Updated: 2023年11月9日

huggingface.co

TheBloke/SOLAR-10.7B-Instruct-v1.0-uncensored-GGUF

Total runs: 12.7K

Run Growth: 1.7K

Growth Rate: 13.03%

Updated: 2023年12月19日

huggingface.co

TheBloke/Noromaid-13B-v0.3-GGUF

Total runs: 11.6K

Run Growth: -7.5K

Growth Rate: -60.88%

Updated: 2024年1月7日

huggingface.co

TheBloke/Luna-AI-Llama2-Uncensored-GGUF

Total runs: 11.3K

Run Growth: 2.5K

Growth Rate: 21.79%

Updated: 2023年9月27日

huggingface.co

TheBloke/OpenHermes-2.5-Mistral-7B-GGUF

Total runs: 11.1K

Run Growth: 5.7K

Growth Rate: 51.37%

Updated: 2023年11月2日

huggingface.co

TheBloke/dolphin-2.7-mixtral-8x7b-GGUF

Total runs: 10.8K

Run Growth: 1.5K

Growth Rate: 14.00%

Updated: 2024年1月1日

huggingface.co

TheBloke/Llama-2-13B-chat-GGUF

Total runs: 9.0K

Run Growth: 0

Growth Rate: 0.00%

Updated: 2023年9月27日

huggingface.co

TheBloke/CapybaraHermes-2.5-Mistral-7B-GGUF

Total runs: 8.9K

Run Growth: 1.7K

Growth Rate: 19.77%

Updated: 2024年1月31日

huggingface.co

TheBloke/openchat_3.5-AWQ

Total runs: 8.8K

Run Growth: 8.6K

Growth Rate: 97.47%

Updated: 2023年11月9日

huggingface.co

TheBloke/Emerhyst-20B-GGUF

Total runs: 8.7K

Run Growth: 7.9K

Growth Rate: 90.39%

Updated: 2023年9月28日

huggingface.co

TheBloke/CodeLlama-7B-GGUF

Total runs: 8.7K

Run Growth: 2.4K

Growth Rate: 28.68%

Updated: 2023年9月27日

huggingface.co

TheBloke/LlamaGuard-7B-AWQ

Total runs: 8.6K

Run Growth: 3.4K

Growth Rate: 45.96%

Updated: 2023年12月11日

huggingface.co

TheBloke/Llama-2-7B-fp16

Total runs: 8.4K

Run Growth: 1.4K

Growth Rate: 23.53%

Updated: 2023年8月27日

huggingface.co

TheBloke/dolphin-2.5-mixtral-8x7b-GGUF

Total runs: 8.3K

Run Growth: 1.9K

Growth Rate: 25.52%

Updated: 2023年12月14日

huggingface.co

TheBloke/TinyLlama-1.1B-Chat-v0.3-AWQ

Total runs: 8.2K

Run Growth: 2.5K

Growth Rate: 31.78%

Updated: 2023年11月9日

huggingface.co

TheBloke/CausalLM-7B-GGUF

Total runs: 7.7K

Run Growth: 2.8K

Growth Rate: 36.88%

Updated: 2023年10月23日

huggingface.co

TheBloke/TinyLlama-1.1B-Chat-v0.3-GGUF

Total runs: 7.7K

Run Growth: 23

Growth Rate: 0.31%

Updated: 2023年10月3日

huggingface.co

TheBloke/rocket-3B-GGUF

Total runs: 7.7K

Run Growth: 3.3K

Growth Rate: 42.01%

Updated: 2023年11月23日

huggingface.co

TheBloke/Llama-2-7B-Chat-AWQ

Total runs: 7.5K

Run Growth: 3.7K

Growth Rate: 50.44%

Updated: 2023年11月9日

huggingface.co

TheBloke/Mistral-7B-OpenOrca-GGUF

Total runs: 7.4K

Run Growth: 3.6K

Growth Rate: 50.51%

Updated: 2023年10月2日

huggingface.co

TheBloke/MythoMax-L2-Kimiko-v2-13B-GGUF

Total runs: 6.9K

Run Growth: 1.6K

Growth Rate: 22.59%

Updated: 2023年9月27日

huggingface.co

TheBloke/Llama-2-13B-fp16

Total runs: 6.9K

Run Growth: -231

Growth Rate: -3.58%

Updated: 2023年7月20日

huggingface.co

TheBloke/Llama-2-70B-Chat-AWQ

Total runs: 6.8K

Run Growth: 3.4K

Growth Rate: 49.27%

Updated: 2023年11月9日

huggingface.co

TheBloke/deepseek-coder-6.7B-base-AWQ

Total runs: 6.8K

Run Growth: -3.1K

Growth Rate: -45.67%

Updated: 2023年11月9日

huggingface.co

TheBloke/Mixtral-8x7B-v0.1-GGUF

Total runs: 6.7K

Run Growth: 3.1K

Growth Rate: 45.94%

Updated: 2023年12月14日

huggingface.co

TheBloke/Llama-2-13B-chat-AWQ

Total runs: 6.7K

Run Growth: 2.4K

Growth Rate: 35.89%

Updated: 2023年11月9日

huggingface.co

TheBloke/Utopia-13B-GGUF

Total runs: 6.6K

Run Growth: 6.3K

Growth Rate: 94.69%

Updated: 2023年11月3日

huggingface.co

TheBloke/deepseek-coder-6.7B-base-GGUF

Total runs: 6.6K

Run Growth: 4.0K

Growth Rate: 60.80%

Updated: 2023年11月5日

huggingface.co

TheBloke/Llama-2-70B-Chat-GGUF

Total runs: 6.6K

Run Growth: 0

Growth Rate: 0.00%

Updated: 2023年11月21日

huggingface.co

TheBloke/CodeLlama-13B-GGUF

Total runs: 6.5K

Run Growth: 2.8K

Growth Rate: 42.91%

Updated: 2023年9月27日

huggingface.co

TheBloke/deepseek-llm-7B-base-GGUF

Total runs: 6.5K

Run Growth: 6.3K

Growth Rate: 97.13%

Updated: 2023年11月29日

huggingface.co

TheBloke/Llama-2-70B-Chat-GPTQ

Total runs: 6.5K

Run Growth: 211

Growth Rate: 4.12%

Updated: 2023年9月27日

huggingface.co

TheBloke/WizardLM-1.0-Uncensored-Llama2-13B-GGUF

Total runs: 6.4K

Run Growth: 184

Growth Rate: 2.89%

Updated: 2023年9月27日

huggingface.co

TheBloke/Wizard-Vicuna-30B-Uncensored-GGUF

Total runs: 6.3K

Run Growth: -288

Growth Rate: -4.55%

Updated: 2023年9月27日

huggingface.co

TheBloke/Silicon-Maid-7B-GGUF

Total runs: 6.1K

Run Growth: 4.4K

Growth Rate: 70.76%

Updated: 2023年12月27日

huggingface.co

TheBloke/dolphin-2.7-mixtral-8x7b-AWQ

Total runs: 6.1K

Run Growth: -4.4K

Growth Rate: -72.87%

Updated: 2024年1月1日

huggingface.co

TheBloke/CodeLlama-7B-Instruct-GGUF

Total runs: 6.1K

Run Growth: 548

Growth Rate: 9.09%

Updated: 2023年9月27日

huggingface.co

TheBloke/airoboros-mistral2.2-7B-GGUF

Total runs: 6.1K

Run Growth: 3.1K

Growth Rate: 52.77%

Updated: 2023年10月3日

huggingface.co

TheBloke/CodeLlama-7B-Python-GGUF

Total runs: 6.0K

Run Growth: 1.1K

Growth Rate: 18.49%

Updated: 2023年9月27日

huggingface.co

TheBloke/Open_Gpt4_8x7B_v0.2-GGUF

Total runs: 5.8K

Run Growth: 0

Growth Rate: 0.00%

Updated: 2024年1月12日

huggingface.co

TheBloke/meditron-7B-AWQ

Total runs: 5.8K

Run Growth: -77.6K

Growth Rate: -2322.53%

Updated: 2023年11月30日

huggingface.co

TheBloke/Rose-20B-GGUF

Total runs: 5.6K

Run Growth: 5.1K

Growth Rate: 93.19%

Updated: 2023年11月24日

huggingface.co

TheBloke/CodeLlama-13B-Instruct-GGUF

Total runs: 5.5K

Run Growth: 951

Growth Rate: 17.26%

Updated: 2023年9月27日

huggingface.co

TheBloke/dolphin-2.2.1-mistral-7B-GGUF

Total runs: 5.3K

Run Growth: 2.1K

Growth Rate: 39.17%

Updated: 2023年10月30日

huggingface.co

TheBloke/Open_Gpt4_8x7B-GGUF

Total runs: 5.0K

Run Growth: 53

Growth Rate: 1.06%

Updated: 2024年1月5日

huggingface.co

TheBloke/Wizard-Vicuna-7B-Uncensored-GGUF

Total runs: 5.0K

Run Growth: -521

Growth Rate: -10.25%

Updated: 2023年9月27日

huggingface.co

TheBloke/zephyr-7B-beta-AWQ

Total runs: 5.0K

Run Growth: 1.4K

Growth Rate: 29.08%

Updated: 2023年11月9日

huggingface.co

TheBloke/OpenHermes-2.5-Mistral-7B-AWQ

Total runs: 4.9K

Run Growth: 1.9K

Growth Rate: 39.79%

Updated: 2023年11月9日

huggingface.co

TheBloke/llama2_70b_chat_uncensored-GGUF

Total runs: 4.9K

Run Growth: -967

Growth Rate: -19.41%

Updated: 2023年9月27日

huggingface.co

TheBloke/dolphin-2.6-mistral-7B-GGUF

Total runs: 4.9K

Run Growth: 2.1K

Growth Rate: 43.40%

Updated: 2023年12月28日

huggingface.co

TheBloke/TinyLlama-1.1B-Chat-v0.3-GPTQ

Total runs: 4.7K

Run Growth: 0

Growth Rate: 0.00%

Updated: 2023年10月3日

huggingface.co

TheBloke/claude2-alpaca-13B-GGUF

Total runs: 4.7K

Run Growth: 2.8K

Growth Rate: 60.03%

Updated: 2023年11月10日

huggingface.co

TheBloke/Nous-Hermes-2-Mixtral-8x7B-DPO-GPTQ

Total runs: 4.6K

Run Growth: 2.2K

Growth Rate: 89.73%

Updated: 2024年1月16日

huggingface.co

TheBloke/wizardLM-7B-HF

Total runs: 4.5K

Run Growth: 3.3K

Growth Rate: 71.37%

Updated: 2023年6月5日

huggingface.co

TheBloke/deepseek-coder-1.3b-base-GGUF

Total runs: 4.4K

Run Growth: 3.9K

Growth Rate: 89.89%

Updated: 2023年11月5日

huggingface.co

TheBloke/Mistral-7B-Instruct-v0.1-GPTQ

Total runs: 4.4K

Run Growth: 1.3K

Growth Rate: 30.90%

Updated: 2023年9月29日

huggingface.co

TheBloke/em_german_mistral_v01-GGUF

Total runs: 4.3K

Run Growth: 167

Growth Rate: 4.82%

Updated: 2023年10月10日

huggingface.co

TheBloke/koala-13B-HF

Total runs: 4.2K

Run Growth: 2.4K

Growth Rate: 58.91%

Updated: 2023年6月5日

huggingface.co

TheBloke/Wizard-Vicuna-13B-Uncensored-GPTQ

Total runs: 4.2K

Run Growth: 240

Growth Rate: 5.81%

Updated: 2023年9月27日

huggingface.co

TheBloke/dolphin-2.1-mistral-7B-GGUF

Total runs: 4.1K

Run Growth: 1.4K

Growth Rate: 35.26%

Updated: 2023年10月22日

huggingface.co

TheBloke/NeuralBeagle14-7B-GPTQ

Total runs: 4.1K

Run Growth: -7.2K

Growth Rate: -116.41%

Updated: 2024年1月17日

huggingface.co

TheBloke/Wizard-Vicuna-7B-Uncensored-GPTQ

Total runs: 4.0K

Run Growth: 275

Growth Rate: 6.99%

Updated: 2023年9月27日

huggingface.co

TheBloke/Mistral-7B-Claude-Chat-GGUF

Total runs: 4.0K

Run Growth: 1.7K

Growth Rate: 42.49%

Updated: 2023年10月28日

huggingface.co

TheBloke/koala-7B-HF

Total runs: 4.0K

Run Growth: 3.1K

Growth Rate: 80.33%

Updated: 2023年6月5日

huggingface.co

TheBloke/WhiteRabbitNeo-13B-GGUF

Total runs: 3.8K

Run Growth: 2.5K

Growth Rate: 65.53%

Updated: 2023年12月21日

huggingface.co

TheBloke/zephyr-7B-beta-GPTQ

Total runs: 3.8K

Run Growth: 1.4K

Growth Rate: 37.37%

Updated: 2023年10月27日

huggingface.co

TheBloke/Yarn-Mistral-7B-128k-GGUF

Total runs: 3.6K

Run Growth: 1.1K

Growth Rate: 32.56%

Updated: 2023年11月2日

huggingface.co

TheBloke/Phind-CodeLlama-34B-v2-GGUF

Total runs: 3.6K

Run Growth: 739

Growth Rate: 20.50%

Updated: 2023年9月27日

huggingface.co

TheBloke/WizardCoder-Python-7B-V1.0-GGUF

Total runs: 3.6K

Run Growth: 2.0K

Growth Rate: 58.21%

Updated: 2023年9月27日

huggingface.co

TheBloke/CodeLlama-34B-GGUF

Total runs: 3.5K

Run Growth: -332

Growth Rate: -9.96%

Updated: 2023年9月27日

huggingface.co

TheBloke/WizardLM-13B-V1.2-GGUF

Total runs: 3.4K

Run Growth: 612

Growth Rate: 17.80%

Updated: 2023年9月27日

TheBloke / Rogue-Rose-103b-v0.2-AWQ

Introduction of Rogue-Rose-103b-v0.2-AWQ

Model Details of Rogue-Rose-103b-v0.2-AWQ

Rogue Rose 103B v0.2 - AWQ

Description

About AWQ

Repositories available

Prompt template: Vicuna-Short

Provided files, and AWQ parameters

How to easily download and use this model in text-generation-webui

Multi-user inference server: vLLM

Multi-user inference server: Hugging Face Text Generation Inference (TGI)

Inference from Python code using Transformers

Install the necessary packages

Transformers example code (requires Transformers 4.35.0 and later)

Compatibility

Discord

Thanks, and how to contribute

Original model card: Sophosympatheia's Rogue Rose 103B v0.2

Overview

Sampler Tips

Prompt Tips

Quantizations

Licence and usage restrictions

Tools Used

Runs of TheBloke Rogue-Rose-103b-v0.2-AWQ on huggingface.co

More Information About Rogue-Rose-103b-v0.2-AWQ huggingface.co Model

More Rogue-Rose-103b-v0.2-AWQ license Visit here:

Rogue-Rose-103b-v0.2-AWQ huggingface.co

Rogue-Rose-103b-v0.2-AWQ huggingface.co Url

TheBloke Rogue-Rose-103b-v0.2-AWQ online free

TheBloke Rogue-Rose-103b-v0.2-AWQ online free url in huggingface.co:

Rogue-Rose-103b-v0.2-AWQ install

Rogue-Rose-103b-v0.2-AWQ install url in huggingface.co:

Url of Rogue-Rose-103b-v0.2-AWQ

Rogue-Rose-103b-v0.2-AWQ huggingface.co Url

Provider of Rogue-Rose-103b-v0.2-AWQ huggingface.co

Other API from TheBloke