Activity Feed

AI & ML interests

None defined yet.

Recent Activity

Articles

amd 's collections 55

HyLo: Long-Context Aware Upcycling for Hybrid LLMs
Hybrid MLA + Gated DeltaNet / Mamba-2 models from long-context aware upcycling: up to 32x usable context with over 90% KV-cache reduction.
Instella-MoE ✨
Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs.
zentorch TorchAO Quantized Models - PyTorch 2.11
TorchAO v0.17.0 quantized models for AMD EPYC CPU inference.
RyzenAI-1.3_LLM_Hybrid_Models
Models quantized by Quark and prepared for the OGA-based hybrid execution flow (Ryzen AI 1.3)
zentorch TorchAO Quantized Models - PyTorch 2.10
TorchAO quantized models for AMD EPYC CPU inference. The inference stack includes vLLM (0.15.0 to 0.18.0), PyTorch 2.10, and zentorch 5.2.1.
RyzenAI-1.3_LLM_NPU_Models
Models quantized by Quark and prepared for the OGA-based NPU-only execution flow (Ryzen AI 1.3)
Quark Quantized ONNX LLMs for Ryzen AI 1.3 EA
ONNX Runtime generate() API based models quantized by Quark and optimized for Ryzen AI Strix Point NPU
HyLo: Long-Context Aware Upcycling for Hybrid LLMs
Hybrid MLA + Gated DeltaNet / Mamba-2 models from long-context aware upcycling: up to 32x usable context with over 90% KV-cache reduction.
Instella-MoE ✨
Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs.
zentorch TorchAO Quantized Models - PyTorch 2.11
TorchAO v0.17.0 quantized models for AMD EPYC CPU inference.
zentorch TorchAO Quantized Models - PyTorch 2.10
TorchAO quantized models for AMD EPYC CPU inference. The inference stack includes vLLM (0.15.0 to 0.18.0), PyTorch 2.10, and zentorch 5.2.1.
RyzenAI-1.3_LLM_NPU_Models
Models quantized by Quark and prepared for the OGA-based NPU-only execution flow (Ryzen AI 1.3)
RyzenAI-1.3_LLM_Hybrid_Models
Models quantized by Quark and prepared for the OGA-based hybrid execution flow (Ryzen AI 1.3)
Quark Quantized ONNX LLMs for Ryzen AI 1.3 EA
ONNX Runtime generate() API based models quantized by Quark and optimized for Ryzen AI Strix Point NPU