flash-attention
Speed up long-sequence transformer training and inference.
Install
npx skills add NousResearch/hermes-agent@flash-attentionRuns in your terminal. Adds the skill globally for Claude Code, Cursor, Codex and others; add -g -y to skip prompts.
What it does
Flash Attention - Fast Memory-Efficient Attention Quick start Flash Attention provides 2-4x speedup and 10-20x memory reduction for transformer attention through IO-aware tiling and recomputation. PyTorch native (easiest, PyTorch 2.2+): flash-attn library (more features): Common workflows Workflow 1: Enable in existing PyTorch model Copy this checklist: Step 1: Check PyTorch version If 512 tokens. Step 4: Test accuracy matches baseline Workflow 2: Use flash-attn library for advanced features For multi-query attention, sliding window, or H100 FP8. Copy this checklist: Step 1: Install…
Excerpt from the skill's own SKILL.md. Read the full file on GitHub before installing: skills run with your agent's permissions.
View source on GitHubBefore you install
Skills are plain text instructions the agent follows, sometimes with scripts. Check the source, prefer repositories with many installs and stars, and read any script it ships.
Categories
Related skills
- ascii-art
NousResearch/hermes-agent
ASCII art: pyfiglet, cowsay, boxes, image-to-ascii.
412 installs 240K - comfyui
NousResearch/hermes-agent
Generate images, video, and audio via diffusion workflows.
382 installs 240KDeploy & DevOpsDocumentationProductivity - touchdesigner-mcp
NousResearch/hermes-agent
Control TouchDesigner via twozero MCP.
316 installs 240KDataAI & agents - baoyu-comic
NousResearch/hermes-agent
Knowledge comics (知识漫画): educational, biography, tutorial.
92 installs 240KDesign & UIAI & agentsWriting - baoyu-article-illustrator
NousResearch/hermes-agent
Article illustrations: type × style × palette consistency.
89 installs 240KDesign & UIAI & agentsWriting - minecraft-modpack-server
NousResearch/hermes-agent
Host modded Minecraft servers (CurseForge, Modrinth).
88 installs 240KSecurity