Reducing a developer's performance to a single number, without explaining where it comes from, is unfair — and useless for management. That is why the DevScore is made up of six dimensions, each answering a different question about the work. The score from 0 to 100 shows the overview; the dimensions explain how it is composed.
The six dimensions fall into two families: automatic signals, collected from the workflow that already exists (Git, time tracking, AI tools), and leadership reviews, done by the team's manager and tech lead. Together, they reduce unfair readings: a number without context misleads, and so does perception without a number.
Automatic signals
- Delivery pace: commits per business day and pull requests per week. Measures how often technical work turns into reviewable delivery — a healthy pace reduces the risk of surprises at the end of the cycle.
- Code contribution: alive lines (lines that remain active in the product) and change entropy. Measures contribution that leaves a material mark, including useful removals during refactoring — not volume for volume's sake.
- AI adoption: tokens consumed per week in the tools the company provides. Measures actual usage, not value delivered — which is why it should be read alongside the others.
- Logged hours: hours logged per week. Without reliable time tracking, capacity and cost become opinion; with it, planning and predictability have a foundation.
Leadership reviews
- Hard skills (scale of 1 to 9): code quality, architecture, best practices, and command of the technologies, assessed by the technical lead. Changes slowly — it describes maturity, not the mood of the period.
- Soft skills (scale of 1 to 6): communication, collaboration, accountability, and organization. Measures impact on how the team functions, not likability. Self-assessment is not part of the calculation.
Weights that follow the role
No dimension has a fixed universal weight. The weights are configurable and always add up to 100%, reflecting what the company values at each moment and for each role. The expectation for a junior is not the same as for a senior: it is common to give more weight to pace and hours early in a career (building cadence) and more weight to contribution and hard skills at senior levels (depth instead of volume).
Reference ranges and the ceiling
Each signal is compared against a healthy range of operation — with a floor and a ceiling — and not against the best person on the team. Below the range, the signal is insufficient. Within it, the score grows gradually. Above it, the score saturates: more commits, more hours, or more tokens do not buy score. The ranges are also configurable by level, contract type, and nature of the work, with governance: adjustments apply only to upcoming closings.
How to read the dimensions in practice
Two people can end up with the same final score for completely different reasons. The right way to read it is always through the composition: where there is consistency, where there is room to grow. And a low score is a starting point for investigation, never a verdict — it may be a blocker, missing data, or work that automatic metrics capture poorly.
Want to see the details of each dimension, with example ranges and the logic behind weights by seniority? Visit the Devint methodology page.
