Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
Abstract
Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user. Developers share skills by copying them between repositories, which makes them a software supply chain without a registry, versions or provenance. The origin of a copied skill, the reach of a security fix and the repositories that warrant review are therefore unknown. Studies that record which repositories hold a skill at a single point in time cannot reveal who copied it from whom. We contribute the first dated copy network of agent skills, built from the git history of every SKILL.md in GitSkills and covering 2,193,119 skill adoptions across GitHub, together with an interactive viewer. A few repositories are the source of almost all copies, and GitHub stars do not identify them. Skill copies almost never change with their source, and a fix at the source therefore rarely reaches them. We fit a model of which repositories others copy from and use it to rank repositories for audit. Reviewing the 100 repositories it ranks highest prevents 14.9% of later adoptions of high-risk skills, against 0.5% for the 100 most starred, which gives security engineers a short list to check before a skill spreads. Platforms should therefore distribute versioned references rather than copies. Project Website: https://fahdseddik.github.io/Skill-Constellations/
Community
Agent skills are SKILL.md instructions and scripts that AI coding agents such as Claude Code and Codex run with the permissions of their user. Developers share skills by copying them between repositories, which makes them a software supply chain without a registry, versions or provenance. The origin of a copied skill, the reach of a security fix and the repositories that warrant review are therefore unknown. Studies that record which repositories hold a skill at a single point in time cannot reveal who copied it from whom. We contribute the first dated copy network of agent skills, built from the git history of every SKILL.md in GitSkills and covering 2,193,119 skill adoptions across GitHub, together with an interactive viewer. A few repositories are the source of almost all copies, and GitHub stars do not identify them. Skill copies almost never change with their source, and a fix at the source therefore rarely reaches them. We fit a model of which repositories others copy from and use it to rank repositories for audit. Reviewing the 100 repositories it ranks highest prevents 14.9% of later adoptions of high-risk skills, against 0.5% for the 100 most starred, which gives security engineers a short list to check before a skill spreads. Platforms should therefore distribute versioned references rather than copies.
My skills directory is half gists and half random repos, and I genuinely couldn't tell you which commit any of them came from — so the provenance problem is real, and worth mapping. Where I'd push back is the copy network itself: it's built on content similarity, and one edited line breaks the fingerprint. Which means the skills that actually got maintained — the ones a security fix would need to reach — are exactly the ones that have diverged past the threshold, so the graph can't see them. You end up overcounting verbatim copies and undercounting the live, edited subset that matters. I'd want the same graph re-run with fuzzy matching or an AST-level diff over the bundled scripts, plus a histogram of how many copies differ by fewer than five lines. If that histogram is fat, the "reach of a fix" number is measuring the wrong population.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents (2026)
- Agent Skill Evolution: How Revisions Affect Coding Agents (2026)
- Scanning the Harness: Configuration Exposures in AI Coding-Agent Supply Chains (2026)
- A Few Pages of Markdown: Committed AI Configuration and Lower Quality Cost after Coding-Agent Adoption (2026)
- Loop Engineering: Building Blocks, Adoption, and Impact (2026)
- Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents (2026)
- Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.11169 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper