AI Rank
Open SourceModelsAppsInsightArticles
Get the app
← Developers
#29067 108

Yucheng Li

@liyucheng09
Open on GitHub
383
Weighted score
51
Contributions
5
Repos
Top repos
  • microsoft/MInference
  • xlite-dev/Awesome-LLM-Inference
  • microsoft/LLMLingua
  • open-compass/opencompass
  • Zefan-Cai/KVCache-Factory
Ranked AI repos2
19095018
xlite-dev/Awesome-LLM-Inference
+4today

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

5.53k· Python· Lists
10436251
microsoft/MInference
+0today

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

1.23k· Python· Infrastructure
AI Rank

Your AI radar. Open source, models and apps, ranked every day from real momentum and real usage.

Get AI Rank for iPhone
Explore
  • Open Source
  • Models
  • Apps
  • Insight
  • Articles
About
  • About AI Rank
  • Content & rights
  • Terms of use
  • Sources & methodology
  • Top rated
  • Support
  • Privacy
  • iOS app
Data: GitHub; model and app usage from OpenRouter (CC BY 4.0); model ratings from Arena (CC BY 4.0). Provider logos from Lobe Icons (MIT).
© 2026 AI Rank
Open SourceModelsAppsInsight