Measured skill scorecards
Every scorecard below comes from running the skill against a fixed task set, twice — once with the skill's guidance available, once without — and judging both runs with a model that did not do the work, against a rubric with fixed dimensions for failure-mode awareness, actionable specificity, and a high-risk blacklist. This is a measurement of output quality, not a marketing number: negative results are published here too.
Content
| Pack | Skill evaluated | n | Executor / Judge | Pass: no-skill → with-skill | Δpp | Rubric (Δ) | Preferred | Composite | Date |
|---|---|---|---|---|---|---|---|---|---|
| humanizer | humanizer | 3 | claude/sonnet · gemini | 47% → 43% | -3.3 [-16.7, 6.7] | 78.8 (+9.0) | 12/30 noise | 43 | 9/14/2026 |
Findings & example
Rewrite the text in inputs.text so it sounds like someone wrote it naturally. Keep everything it says. Inputs: { "text": "Additionally, our new dashboard represents a pivotal step forward. The team meticulously crafted every screen, and the results underscore our commitment to quality." } With skill The new dashboard is a real step forward. The team designed every screen carefully, and it shows. Without skill Also, our new dashboard is a big step forward. The team put real care into every screen, and it shows—the quality speaks for itself. | |||||||||
Real Estate
| Pack | Skill evaluated | n | Executor / Judge | Pass: no-skill → with-skill | Δpp | Rubric (Δ) | Preferred | Composite | Date |
|---|---|---|---|---|---|---|---|---|---|
| realtor_listing_copy | realtor-listing-copy | 3 | claude/sonnet · gemini | 10% → 93% | +83.3 [66.7, 96.7] | 95.3 (+43.1) | 29/30 gain | 93 | 9/14/2026 |
Findings & example
Write listing copy (headline + body) for the property in inputs. Inputs: { "facts": { "address": "88 Sycamore Ridge", "price": "$459,000", "beds_baths": "3 bed / 2.5 bath", "sqft": 1920, "year_built": 2010, "features": [ "renovated primary bathroom (2021)", "open-concept floor plan", "attached two-car garage", "covered back patio" ], "nearby": "0.4 miles from Maple Elementary School" } } With skill ``` Headline: Open-Concept 3-Bed on Sycamore Ridge With a 2021 Primary Bath Remodel 88 Sycamore Ridge offers 1,920 square feet of open-concept living, anchored by a primary bathroom renovated in 2021. The layout provides an easy flow between living spaces. Three bedrooms and two and a half bathrooms, with an attached two-car garage and a covered back patio. Built in 2010. Maple Elementary School is 0.4 miles away. Key facts: - Beds/Baths: 3 / 2.5 - Square footage: 1,920 sq ft - Year built: 2010 - Price: $459,000 Without skill **Charming Family Home Just Steps from Maple Elementary — 88 Sycamore Ridge** Welcome to 88 Sycamore Ridge, a beautifully maintained 3-bedroom, 2.5-bath home offering 1,920 square feet of thoughtfully designed living space. Built in 2010, this residence features an open-concept floor plan that seamlessly connects the kitchen, dining, and living areas — perfect for both everyday living and entertaining. The primary bathroom was fully renovated in 2021, offering a fresh, modern retreat at the end of the day. Step outside to enjoy the covered back patio, an ideal spot for morning coffee or evening relaxation, and take advantage of the attached two-car garage for convenient storage and parking. Located just 0.4 miles from Maple Elementary School, this home is perfectly positioned for families seeking both comfort and convenience. Offered at $459,000 — schedule your showing today! | |||||||||
Uncategorized
| Pack | Skill evaluated | n | Executor / Judge | Pass: no-skill → with-skill | Δpp | Rubric (Δ) | Preferred | Composite | Date |
|---|---|---|---|---|---|---|---|---|---|
| linkedin_skills | linkedin-post-writer | 3 | claude/sonnet · gemini | 57% → 30% | -26.7 [-53.3, 3.3] | 68.1 (-11.7) | 3/30 loss | 30 | 9/14/2026 |
Findings & example
Write a LinkedIn post about the topic in inputs. First line must be under 210 characters (it's the preview hook). Use 3 or fewer hashtags. Inputs: { "topic": "we just shipped a feature that cut our support ticket volume by 30% after 6 weeks of user interviews", "goal": "comments" } With skill 30% drop in support tickets. One change: six weeks of user interviews before we wrote a line of code. We almost shipped the wrong fix. Our first instinct was to add more tooltips. Cleaner docs, better onboarding copy, a redesigned help center. The usual playbook. Then we sat down with 14 users over six weeks and asked them to walk us through the exact moment they opened a ticket. Turns out the problem wasn't that people couldn't find help. It was that the export button lived three menus deep, and by the time someone found it, they'd already given up and emailed support instead. We moved one button. Renamed it to match the word users actually said out loud in every single call: "download," not "export." Ticket volume dropped 30% over the following six weeks. No new docs. No AI chatbot. No redesign. The interviews cost us six weeks we didn't think we had. The wrong fix would have cost us a quarter, and we'd have still had the tickets. What's the last support ticket your team fixed with a feature, instead of a five-minute conversation with the person filing it? P.S. Happy to share the exact interview script we used if it's useful. #ProductManagement #UserResearch #SaaS Without skill 6 weeks of user interviews. 1 feature. 30% fewer support tickets. Here's what most teams get wrong: they build first, then wonder why adoption is low and tickets keep piling up. We flipped it. For six weeks, our team sat down with users, watched them struggle, asked "why" one too many times, and resisted the urge to jump straight to solutions. The pattern that emerged wasn't what we expected. Users weren't confused about how to use the product — they were confused about what would happen before they clicked anything. So we built clarity into the interface itself. No new docs. No extra onboarding steps. Just answers, upfront, exactly where the questions were happening. Result: a 30% drop in support tickets in the first six weeks post-launch. The lesson? Sometimes the biggest wins aren't new features — they're the right questions asked long enough to hear the real problem. Curious — what's a "small" change your team made that had an outsized impact on support or retention? Would love to hear your stories below. 👇 #ProductManagement #UserResearch #CustomerSuccess | |||||||||
Methodology
Each pack's eval suite runs the same task set through two arms — with the skill's guidance loaded, and without — at a fixed, pinned model profile ("lean-v1"). Only the final deliverable is graded: drafts, notes, and other scaffolding are stripped before any check or judge call runs, and whether the output was paste-ready is reported alongside the score rather than folded into it. Deterministic verifiers check what they can (character limits, required content, banned phrases, numbers grounded in the task's own facts); an independent judge model — never the same model that produced the output, and with its rubric criteria shown in a freshly shuffled order on every call to avoid position bias — scores the rest against three mandatory dimensions (failure-mechanism awareness, actionable specificity, a high-risk-action blacklist) plus pack-specific items. Separately, a blind paired comparison shows the judge both arms' deliverables for the same task with no label saying which is which, in both orders; a side only wins if it's preferred in both orders; a split decision is scored as a tie, not a coin flip. That paired comparison, not the raw score gap, is the primary verdict (gain / loss / noise) shown on each scorecard. Pass rates and score deltas carry Wilson or paired-bootstrap 95% confidence intervals; a pack's headroom (how much score the no-skill arm left on the table) and any low-signal tasks (where the no-skill arm already scores too high for a skill to show a gain) are called out on the scorecard, alongside the token-cost ratio between the two arms. A task passes only above an 80% rubric-item threshold; a re-run must beat the currently accepted score by at least 8 percentage points (a threshold derived from measured run-to-run and judge-to-judge spread, not picked arbitrarily) before it replaces it.
This page shows only packs that have been re-measured under this method — a pack still on an older, single-score result is held back until it's re-run.
What this doesn't measure
This is a controlled measurement against a fixed task set, not a guarantee of real-world outcomes. Composite and rubric score can disagree in sign on a small task set — a strict pass/fail collapse is brittle at n=3 with 8 tasks, so both numbers are shown. Two independently valid judge models can disagree on the same output by several points; a re-run under a different judge is never compared directly against one under another. For a pack that bundles many sub-skills, the scorecard reflects only the sampled sub-skill named in the table, not the whole pack. A high no-skill baseline leaves little room for any skill to show a gain — see the headroom figure on a pack's scorecard when it's reported. The rubric is Skills Wiki's own adaptation of SkillLens's three dimensions and has not been validated against human graders. The token-cost ratio excludes judge tokens on backends that don't report their own usage. The paired comparison skips any task with no judge call at all (a purely deterministic check), so its sample size can be smaller than the number of records in each arm.
Sources
Curated, evaluated skills: SkillsBench, arXiv:2602.12670. Rubric-guided selection of skill documents: SkillLens, arXiv:2605.23899.