-
Notifications
You must be signed in to change notification settings - Fork 600
Pull requests: NVIDIA/Model-Optimizer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Add Llama to Puzzletron v2
puzzletron_v2
Related to feature/puzzletron_v2 branch
#2454
opened Sep 17, 2026 by
grzegorz-k-karch
Contributor
•
Draft
Fix hybrid stack spec serialization in Megatron-Bridge checkpoints
#2452
opened Sep 16, 2026 by
kevalmorabia97
Collaborator
•
Draft
[6771663] Preserve ONNX API output types when wiring casts
#2451
opened Sep 16, 2026 by
ajrasane
Contributor
Loading…
[OMNIML-5899] Add IQ post-training quantization recipes
#2449
opened Sep 16, 2026 by
hychiang-git
Contributor
Loading…
[OMNIML-5899] Add CUDA kernels for IQ packing
#2448
opened Sep 16, 2026 by
hychiang-git
Contributor
Loading…
[OMNIML-5899] Export IQ checkpoints from HF and Megatron
#2447
opened Sep 16, 2026 by
hychiang-git
Contributor
Loading…
[OMNIML-5899] Add IQ quantization codecs and backend
#2446
opened Sep 16, 2026 by
hychiang-git
Contributor
Loading…
docs: require evaluation provenance for model-card generation settings
#2441
opened Sep 15, 2026 by
chadvoegele
Contributor
•
Draft
Fix HF export crash when a dynamic-block quantizer has zero amax
cherry-pick-0.47.0
Upcoming release
#2438
opened Sep 15, 2026 by
yueshen2016
Contributor
Loading…
[https://nvbugspro.nvidia.com/bug/6778095] Fix fused P-QDQ to respect disabled quantization during calibration
#2434
opened Sep 15, 2026 by
yingguo-trt
Loading…
Fail fast on non-finite AutoQuantize output gradients
#2432
opened Sep 14, 2026 by
meenchen
Contributor
Loading…
[DRAFT] Make distributed Puzzletron campaigns reliable
#2430
opened Sep 14, 2026 by
chochowski
Contributor
Loading…
Carry unplaced checkpoint weights using the loader's accounting, replacing MTP name-matching
#2427
opened Sep 14, 2026 by
shengliangxu
Collaborator
Loading…
Fixes ONNX validation and AutoCast handling for ONNX Runtime legacy operators
#2425
opened Sep 13, 2026 by
haoxiz-nvidia
Contributor
Loading…
fix(quantization): correct the affine-quant bias config contract
#2422
opened Sep 12, 2026 by
sriharshapy
Loading…
Fix NF4 dequantize when the scales span more than one int8 group
#2421
opened Sep 12, 2026 by
rootkiller6788
Loading…
Fix AutoCast unknown dimension metadata
#2420
opened Sep 12, 2026 by
haoxiz-nvidia
Contributor
Loading…
Env: update onnxruntime and cuda flag
#2419
opened Sep 12, 2026 by
haoxiz-nvidia
Contributor
Loading…
[6466805] Add Dynamo ONNX export support for quantized models
#2418
opened Sep 12, 2026 by
ajrasane
Contributor
Loading…
Previous Next
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.