PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX Paper • 2608.17379 • Published 26 days ago • 4
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX Paper • 2608.17379 • Published 26 days ago • 4
Adaptive Self-improvement LLM Agentic System for ML Library Development Paper • 2502.02534 • Published Sep 19, 2025
CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models Paper • 2404.08763 • Published Apr 12, 2024 • 2
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization Paper • 2511.15915 • Published Apr 15 • 5
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization Paper • 2511.15915 • Published Apr 15 • 5
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity Paper • 2501.16295 • Published Jan 27, 2025 • 9
Learning to (Learn at Test Time): RNNs with Expressive Hidden States Paper • 2407.04620 • Published Jul 5, 2024 • 34
MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression Paper • 2406.14909 • Published Jun 21, 2024 • 16
thrunlab/sparse_llama_7b_hf2_refined_web_90p_2024-05-12 Text Generation • 7B • Updated May 12, 2024 • 9
thrunlab/sparse_llama_7b_hf2_refined_web_50p_2024-05-12 Text Generation • 7B • Updated May 12, 2024 • 18
thrunlab/sparse_mistral_7b_refined_web_50p_2024-05-11 Text Generation • 7B • Updated May 11, 2024 • 11
thrunlab/sparse_llama_7b_hf2_refined_web_50p_2024-05-11 Text Generation • 7B • Updated May 11, 2024 • 11