Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 18 additions & 17 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@

## Latest News

* 07/23/2026 7.3.1 `main`: ✨ Added `inkling_mm_model` model support
* 07/22/2026 7.3.1 `main`: ✨ Added Poolside `Laguna S 2.1` model support
* 07/14/2026 7.3.0-dev `main`: ✨ Added `nemotron_h_puzzle` model support
* 07/07/2026 7.3.0-dev `main`: ✨ Added `deepseek_vl` model support
Expand Down Expand Up @@ -258,24 +259,24 @@ Selected public references where teams or companies explicitly mention GPT-QMode

## Model Support

| Model | | | | | | | | | |
|-------------------------------|---|---------------------------------|--|------------------|--|---------------------------------|--|------------------------|---|
| Apertus | ✅ | EXAONE 3/4 | ✅ | Dots1 | ✅ | Mistral3 / Ministral3 | ✅ | Qwen 2/3/3.5 (Next/MoE) | ✅ |
| Baichuan | ✅ | Falcon (H1 / Mamba) | ✅ | InternLM 1/2/2.5 | ✅ | Mixtral | ✅ | Qwen 2/2.5/3 VL | ✅ |
| Bloom | ✅ | FastVLM | ✅ | Kimi K2 | ✅ | MobileLLM | ✅ | Qwen 2.5/3 Omni | ✅ |
| ChatGLM | ✅ | Gemma 1-4 / 3n | ✅ | Klear | ✅ | MOSS | ✅ | RefinedWeb | ✅ |
| CodeGen | ✅ | GPTBigCode | ✅ | LING/RING | ✅ | MPT | ✅ | StableLM | ✅ |
| Cohere 1-2 / 2 MoE | ✅ | GPT-Neo / NeoX | ✅ | Llama 1-3.3 | ✅ | Nemotron H / H Puzzle / Omni | ✅ | StarCoder2 | ✅ |
| DBRX Converted | ✅ | GPT-2 | ✅ | Llama 3.2 VL | ✅ | Nemotron Ultra / Labs-Diffusion | ✅ | TeleChat2 | ✅ |
| Deci | ✅ | GPT-J | ✅ | Llama 4 | ✅ | OPT | ✅ | Trinity | ✅ |
| DeepSeek-V2/V3/V4/R1 | ✅ | GPT-OSS | ✅ | LongCat Flash | ✅ | OLMo2 / LLaDA2 | ✅ | Yi | ✅ |
| DeepSeek-V2 Lite / VL / VL2 / OCR2 | ✅ | Granite / Granite MoE | ✅ | LongLLaMA | ✅ | Ovis 1.6/2/2.5/2.6 MoE/2.6 Next | ✅ | Seed-OSS | ✅ |
| Dream | ✅ | GRIN-MoE | ✅ | Instella | ✅ | Phi 1-4 | ✅ | Voxtral | ✅ |
| Model | | | | | | | | | |
|-------------------------------|---|---------------------------------|--|----------------------------|--|---------------------------------|--|------------------------|---|
| Apertus | ✅ | EXAONE 3/4 | ✅ | Dots1 | ✅ | Mistral3 / Ministral3 | ✅ | Qwen 2/3/3.5 (Next/MoE) | ✅ |
| Baichuan | ✅ | Falcon (H1 / Mamba) | ✅ | InternLM 1/2/2.5 | ✅ | Mixtral | ✅ | Qwen 2/2.5/3 VL | ✅ |
| Bloom | ✅ | FastVLM | ✅ | Kimi K2 | ✅ | MobileLLM | ✅ | Qwen 2.5/3 Omni | ✅ |
| ChatGLM | ✅ | Gemma 1-4 / 3n | ✅ | Klear | ✅ | MOSS | ✅ | RefinedWeb | ✅ |
| CodeGen | ✅ | GPTBigCode | ✅ | LING/RING | ✅ | MPT | ✅ | StableLM | ✅ |
| Cohere 1-2 / 2 MoE | ✅ | GPT-Neo / NeoX | ✅ | Llama 1-3.3 | ✅ | Nemotron H / H Puzzle / Omni | ✅ | StarCoder2 | ✅ |
| DBRX Converted | ✅ | GPT-2 | ✅ | Llama 3.2 VL | ✅ | Nemotron Ultra / Labs-Diffusion | ✅ | TeleChat2 | ✅ |
| Deci | ✅ | GPT-J | ✅ | Llama 4 | ✅ | OPT | ✅ | Trinity | ✅ |
| DeepSeek-V2/V3/V4/R1 | ✅ | GPT-OSS | ✅ | LongCat Flash | ✅ | OLMo2 / LLaDA2 | ✅ | Yi | ✅ |
| DeepSeek-V2 Lite / VL / VL2 / OCR2 | ✅ | Granite / Granite MoE | ✅ | LongLLaMA | ✅ | Ovis 1.6/2/2.5/2.6 MoE/2.6 Next | ✅ | Seed-OSS | ✅ |
| Dream | ✅ | GRIN-MoE | ✅ | Instella | ✅ | Phi 1-4 | ✅ | Voxtral | ✅ |
| ERNIE 4.5 / MoE / VL MoE | ✅ | GLM 4/4V/4.5V/4.6V/5/5.1/OCR/ASR | ✅ | GLM4 MoE / Lite / 4.5V MoE | ✅ | MiniCPM 3/O/V/V 4_6 | ✅ | PanGu-α | ✅ |
| XVERSE | ✅ | Brumby | ✅ | Hymba | ✅ | Mistral | ✅ | Qwen 1/2/3/3.5 | ✅ |
| MiniMax M2/M3 | ✅ | AfMoE | ✅ | Bailing-MoE | ✅ | LFM2 / LFM2-VL / LFM2-MoE | ✅ | Marin | ✅ |
| InternVL Chat | ✅ | Laguna | ✅ | Mimo / Mimo V2 | ✅ | Zamba / Zamba2 | ✅ | Intern S1 | ✅ |
| HunYuan V1 Dense / MoE | ✅ | HY-V3 | ✅ | | | | | | |
| XVERSE | ✅ | Brumby | ✅ | Hymba | ✅ | Mistral | ✅ | Qwen 1/2/3/3.5 | ✅ |
| MiniMax M2/M3 | ✅ | AfMoE | ✅ | Bailing-MoE | ✅ | LFM2 / LFM2-VL / LFM2-MoE | ✅ | Marin | ✅ |
| InternVL Chat | ✅ | Laguna | ✅ | Mimo / Mimo V2 | ✅ | Zamba / Zamba2 | ✅ | Intern S1 | ✅ |
| HunYuan V1 Dense / MoE | ✅ | HY-V3 | ✅ | Inkling | ✅ | | | | |

Prism Bonsai GGUF checkpoints are supported for inference only through GPT-QModel's native GGUF path and internal GGUF runtime. Bonsai checkpoints load through the normal model path or repo argument and do not require the external `gguf` package. Prism model quantization is not included.

Expand Down
2 changes: 2 additions & 0 deletions gptqmodel/models/auto.py
Original file line number Diff line number Diff line change
Expand Up @@ -125,6 +125,7 @@
from .definitions.hy_v3 import HYV3QModel # noqa: E402
from .definitions.hymba import HymbaQModel # noqa: E402
from .definitions.instella import InstellaQModel # noqa: E402
from .definitions.inkling import InklingMMQModel # noqa: E402
from .definitions.internlm import InternLMQModel # noqa: E402
from .definitions.internlm2 import InternLM2QModel # noqa: E402
from .definitions.interns1 import InternS1QModel # noqa: E402
Expand Down Expand Up @@ -324,6 +325,7 @@
"ovis2_6_next": Ovis2_6_NextQModel,
"telechat": TeleChat2QModel,
"instella": InstellaQModel,
"inkling_mm_model": InklingMMQModel,
"mimo": MimoQModel,
"mimo_v2": MimoV2QModel,
"falcon_h1": FalconH1QModel,
Expand Down
1 change: 1 addition & 0 deletions gptqmodel/models/definitions/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,7 @@
from .hy_v3 import HYV3QModel
from .hymba import HymbaQModel
from .instella import InstellaQModel
from .inkling import InklingMMQModel
from .internlm import InternLMQModel
from .internlm2 import InternLM2QModel
from .interns1 import InternS1QModel
Expand Down
194 changes: 194 additions & 0 deletions gptqmodel/models/definitions/inkling.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,194 @@
# SPDX-FileCopyrightText: 2026 ModelCloud.ai
# SPDX-FileCopyrightText: 2026 qubitium@modelcloud.ai
# SPDX-License-Identifier: Apache-2.0
# Contact: qubitium@modelcloud.ai, x.com/qubitium

from typing import Any, Dict

import torch
from transformers import AutoModelForMultimodalLM, AutoProcessor, ProcessorMixin
from transformers.masking_utils import create_causal_mask, create_sliding_window_causal_mask

from ...utils.calibration import batched
from ...utils.model import MODALITY, move_to
from ...utils.offload import offload_to_disk
from .._const import CPU
from ..base import BaseQModel
from ..moe_lifecycle import GateUpDownMoELifecycleHooks


class InklingMMQModel(BaseQModel):
loader = AutoModelForMultimodalLM

require_load_processor = True
require_trust_remote_code = False
layer_modules_strict = False

modality = [MODALITY.TEXT, MODALITY.IMAGE_TO_TEXT]

dynamic_expert_index = "n_routed_experts"
defuser_auto_detect_moe = True
moe_lifecycle_hooks = GateUpDownMoELifecycleHooks()

pre_lm_head_norm_module = "model.language_model.norm"

# GQA output shapes differ across the attention projections. AWQ should
# optimize the output projection against its actual attention input.
awq_scale_optimize_shape_dependent_modules = ["self_attn.o_proj"]

module_tree = [
"model",
"language_model",
"layers",
"#",
{
"input_layernorm": ("input_layernorm:!",),
"self_attn": (
"q_proj:0",
"k_proj:0",
"v_proj:0",
"r_proj:0",
"q_norm:!",
"k_norm:!",
"o_proj:1",
),
"post_attention_layernorm": ("post_attention_layernorm:!",),
"mlp:moe": {
# Inkling's first decoder blocks use a dense MLP.
"": ("gate_proj:0", "up_proj:0", "down_proj:1"),
"gate": ("gate:!",),
"experts": {
"#": ("gate_proj:0", "up_proj:0", "down_proj:1"),
},
# Shared experts are native 3D parameters rather than Linear
# modules. Keep these two accuracy-sensitive experts dense;
# routed experts are expanded and quantized individually.
},
},
]

def pre_quantize_generate_hook_start(self):
core_model = self.model.model
language_model = core_model.language_model
self.shell_module_materialize(language_model.embed_tokens, self.quantize_config.device)
self.shell_module_materialize(language_model.embed_norm, self.quantize_config.device)
self.shell_module_materialize(core_model.vision_tower, self.quantize_config.device)
self.shell_module_materialize(core_model.audio_tower, self.quantize_config.device)

def pre_quantize_generate_hook_end(self):
core_model = self.model.model
language_model = core_model.language_model
if self.quantize_config.offload_to_disk:
offload_to_disk(
model=language_model,
module=language_model.embed_tokens,
disk_path=self.quantize_config.offload_to_disk_path,
)
offload_to_disk(
model=language_model,
module=language_model.embed_norm,
disk_path=self.quantize_config.offload_to_disk_path,
)
offload_to_disk(
model=core_model,
module=core_model.vision_tower,
disk_path=self.quantize_config.offload_to_disk_path,
)
offload_to_disk(
model=core_model,
module=core_model.audio_tower,
disk_path=self.quantize_config.offload_to_disk_path,
)
return

language_model.embed_tokens = move_to(language_model.embed_tokens, device=CPU)
language_model.embed_norm = move_to(language_model.embed_norm, device=CPU)
core_model.vision_tower = move_to(core_model.vision_tower, device=CPU)
core_model.audio_tower = move_to(core_model.audio_tower, device=CPU)

def load_processor(self) -> ProcessorMixin:
return AutoProcessor.from_pretrained(self.model_local_path, trust_remote_code=False)

@classmethod
def prepare_inputs_for_conversations(
cls,
processor: ProcessorMixin,
conversations: list[dict] | list[list[dict]],
):
return processor.apply_chat_template(
conversations,
tokenize=True,
add_generation_prompt=True,
reasoning_effort="medium",
return_dict=True,
return_tensors="pt",
)

def prepare_dataset(self, calibration_dataset, batch_size: int = 1, **kwargs):
del kwargs
processor = self.load_processor()
calibration_data = []
for batch in batched(calibration_dataset, batch_size):
calibration_data.append(self.prepare_inputs_for_conversations(processor, batch))
del processor
return calibration_data

def move_input_capture_example(
self,
example: Dict[str, Any],
data_device: torch.device,
) -> Dict[str, Any]:
example = super().move_input_capture_example(example, data_device)
pixel_values = example.get("pixel_values")
if not torch.is_tensor(pixel_values) or not pixel_values.is_floating_point():
return example

first_parameter = next(self.model.model.vision_tower.parameters(), None)
if first_parameter is not None:
example["pixel_values"] = pixel_values.to(
device=first_parameter.device,
dtype=first_parameter.dtype,
)
return example

def prepare_layer_replay_kwargs(self, layer, layer_input, additional_inputs, target_device):
additional_inputs = super().prepare_layer_replay_kwargs(
layer,
layer_input,
additional_inputs,
target_device,
)
if not layer_input or not torch.is_tensor(layer_input[0]):
return additional_inputs

self_attn = getattr(layer, "self_attn", None)
layer_config = getattr(self_attn, "config", None)
if layer_config is None:
return additional_inputs

# The first-layer hook captures the sliding layer's prepared 4D mask.
# Rebuild it for every replayed layer so full-attention blocks do not
# accidentally inherit the sliding window. ``conv_mask`` retains the
# original 2D padding signal when the calibration batch is padded.
padding_mask = additional_inputs.get("conv_mask")
if not torch.is_tensor(padding_mask) or padding_mask.ndim != 2:
captured_mask = additional_inputs.get("attention_mask")
padding_mask = captured_mask if torch.is_tensor(captured_mask) and captured_mask.ndim == 2 else None

mask_factory = (
create_sliding_window_causal_mask
if getattr(layer, "layer_type", None) == "hybrid_sliding"
else create_causal_mask
)
additional_inputs["attention_mask"] = mask_factory(
config=layer_config,
inputs_embeds=layer_input[0],
attention_mask=padding_mask,
past_key_values=additional_inputs.get("past_key_values"),
position_ids=additional_inputs.get("position_ids"),
layer_idx=getattr(self_attn, "layer_idx", None),
)
return additional_inputs


__all__ = ["InklingMMQModel"]
Loading