
Learn how LMCache solves the LLM long-context bottleneck by offloading and sharing Key-Value (KV) caches across vLLM engine instances to crush Time-To-First-Token (TTFT).

Learn how FlashInfer delivers high-throughput GPU kernels for LLM serving engines, supporting DeepSeek MLA, FP4/FP8 compute, and hardware from Turing to Blackwell.

Build, trace, evaluate, and optimize production LLM applications with Langfuse, the YC-backed open-source LLM observability platform.

HyperFrames bridges web design and video production, allowing developers and AI agents to transform standard HTML, CSS, and animations into crisp MP4 videos.

Dify turns chaotic LLM development into a unified visual workspace. Build, test, and deploy AI agents and RAG pipelines without rewriting your tech stack.

Break free from vendor lock-in with OpenCodex, a lightweight local proxy that lets you plug DeepSeek, Gemini, Ollama, and Grok right into Claude Code, Codex CLI, and Claude Desktop.

Stop wrestling with image-heavy AI slide generators. PPT Master converts PDFs, docs, and raw ideas into fully editable, native PowerPoint files complete with vector shapes, charts, and voice narration.

Langfuse brings open-source LLM observability, prompt management, evaluations, and cost tracking directly to your AI stack without vendor lock-in.

Take total control over your diffusion models. ComfyUI gives developers and creators a node-based architecture for rapid visual AI execution, memory optimization, and custom API pipelines.

Discover Graft, the open-source context layer that slashes AI coding token costs by 42% and boosts accuracy across Claude Code, Cursor, and Codex.