Engineering Metrics That Don't Suck
The cry for more data and metrics is felt by every leader and every leader's boss. It seems like everyone has been burned by metrics gone wrong, and yet you need to be able to use data to make the right decision and get everyone else on the same page.
The bottom line is that the metric that doesn't suck is the one that actually informs your decisions.
The aim of this page is to demystify engineering metrics so that you can finally have data that is helping you instead of drowning you. Having said that, this is an expansive topic with lots of things to pay attention to.
What Are Engineering Metrics Actually For?
Alright, so there are lots of questions you need answered as a leader. Metrics can help answer those questions. Before we get into that though I want to share a model for how you can go about thinking through the world of metrics so that you can see how to work with them.
Metrics fit into one of four categories: Investigation, Operational, Signals, and Targets.
Investigation Metrics
Simply put, investigation metrics represent data and metrics that don't have real meaning but may reveal an insight when you have a question.
Consider the case of a novel outage in your production systems. You and the teams will likely pore through your investigation data like logging and health metrics to triangulate the cause.
For most organizations and leaders a lot of their dashboards and metric efforts fall into this category. It's data that they hope will reveal an insight, but often it does not.
It's important to note that investigation metrics are important, but because you have to spend so much time searching through them, they are costly to use and other categorizations will likely be more expedient for your needs. So that itch you feel to get more data is likely you hoping if you have more data insights will reveal themselves. That is investigation metrics in action.
Operational Metrics
You know this one, but by the end you are likely going to realize you have work to do. Think of a complex machine with gauges that tell the operator if the machine is working within expected parameters. Those are operational metrics.
Folks use things like DORA, velocity, uptime, test coverage, and all kinds of metrics in this category. Though like all metrics, the intention you have with them matters more than what the metric is.
Operational metrics are something a lot of groups struggle with because while they have the data, they don't have the thresholds that are required for them to be of use.
Consider a thermometer we use when we don't feel well. It shows every single number, but we all know that above 98.6 F you have a fever and above 104 F you're really sick. Those are the thresholds we were taught. A thermometer could simply have a display that says, "Fine," "Fever," and "Emergency!" Instead, we keep the numbers and operate with the well-known thresholds.
Operational metrics need thresholds, otherwise as the numbers shift up and down we can't interpret what they mean.
Signal Metrics
Signals are metrics that act like alarms that prompt you to act. If production goes down you want a signal to alert everyone to stop and investigate. These are tight data-to-action pairs. I call the method I teach to build these, Signal Mapping, and it is a great place to start with building out metrics.
By formulating a "When I see this I will do that" style pair of metric to action you're taking a real stance on meaning and using data to inform actions. It can be scary because you won't always know if the data you picked matches to the action perfectly but with practice you'll be able to do this effortlessly.
Targets
Alright, this is the big one. This is the one that causes lots of dysfunction but I have to address it. Leaders and organizations will set targets in their metrics as goals or indicators of goals being met. At face value this is fine, but things tend to get a little screwy here.
If you're unfamiliar with OKRs, they became hugely popular around 2014 or so and quickly faded away because organizations and leaders struggled to establish key results as targets effectively. The result was that most groups abandoned them and the practice as it became weaponized and sabotaged.
Targets are important and very dangerous when misused. Some advice around target setting that is good to remember is:
- Use relative targets not absolutes (Up by 20% is better than 600 points)
- Expect confusion if you set a target that looks better when it goes down
- Be explicit about the boundaries people can operate within to succeed
- Know when you'll check and what you expect when you check
- Ensure everyone knows what that target means and why it matters
- Balance your target with other metrics
Why Do Engineering Metrics Lie?
Know what everyone trusts? Estimation. I'm kidding. Estimates, the forecasts, the tracking, the status meetings, reports and everything else that comes from wanting to know when something is done are prime examples of lies in our numbers. A team estimates, a leader pads, that leader pads, etc. Deadlines are set in advance of the real deadline because nobody trusts each other, because no estimate was ever right, and manufactured urgency is all we have left. Report after report shows green status even after dates keep slipping, eroding leadership's trust in everything in front of them.
Imagine how much of this activity and theater would be gone if estimates were presented faithfully and nobody had to play these games.
Ever seen teams and leaders game metrics? Of course you have. When teams learn leaders are monitoring velocity, they estimate high so their velocity jumps and they look good even if they deliver the same or less. More on that in a moment.
There's an old saying that you get what you measure and we've all been burned by it. So I want to address that first. If folks know something is being measured or targeted, that number will magically move in the right direction. Unfortunately something will be sacrificed to get there and that's the problem.
Back to the velocity story. I worked in a place this happened and the team gaming their velocity became the favorite of senior leadership. That is, until my team showed up. We didn't play that game, we focused on real results and high quality. The contrast laid the deceit bare and we never looked back. Stupid games get stupid prizes. While the other team inflated estimates to look good, we became the team that shipped consistently with high quality. We were the ones leaders turned to when something had to be on time the first time.
To solve the gaming issue you have to balance your metrics along two dimensions. The first dimension is to pair what you're measuring with the harder version that is likely going to tell a more accurate truth. This is also called leading and lagging indicators but the gist is that not everything that is easy to measure today means what we hope it does in the future. By balancing them along this leading and lagging dimension we can prevent creating and tracking metrics that don't line up to reality. The second dimension to balance along is what is sacrificed. Hint: It's always quality. So how you prevent measures of productivity leading to awful sacrifices in quality is that you also track quality metrics.
This might sound complicated but DORA metrics do this beautifully, so you can look at them as an example of these balancing forces.
I hinted at something a bit ago I want to come back to. Most folks measure what is easy to count. I know that sounds insulting but it's true. Then a magical thing happens where we decide that this easy-to-count thing means this very important other thing. Velocity means productivity. Lines of code or pull-requests mean developer productivity. Test coverage means quality. Easy to count, but those meanings may not be good fits, and if you aren't careful you'll be grossly misled.
When we wind up misusing metrics or when they're lying, they wind up creating a whole lot of extra work for everyone. Folks have to hedge their bets, manage expectations, and work without trust that things will be alright.
Can Developer Productivity Be Measured?
I'm going to give you the brutal truth. Measuring individuals will drive you nuts and your team to quit. There is tons of research on this topic, like Robert Austin's Measuring and Managing Performance in Organizations, showing how most efforts to measure individual productivity work against you.
Here's what I teach leaders. You need to know that your outcomes and results are where they need to be which means looking at a team and system level. Teams are far more than the sum of their parts so observing that is a better indicator for productivity. Second, a poor system or process will destroy everything from your best team to your best developer.
Finally, you do need to know if you have people that are bad apples, are bringing people down, or need help. That's what your 1:1s are for. If you feel compelled to measure individuals, make sure you'd be comfortable with someone measuring you similarly.
Which Metrics Matter for a Software Team?
I'm of two minds about this. On the one hand, I would tell you to start small with one metric and grow from there as your comfort and skill increases. On the other hand, pick an area that needs extra focus and attention even if it is a tricky issue.
The main thing though is that you want to know what meaning that metric has and most importantly, how you'll use it. Is it an operational metric that you need to set thresholds to? Is it a signal that as soon as it goes off you act? You have to know the intent and usage of your metrics.
Most writing on engineering metrics gives you a laundry list of metrics that are fine and not worth me going over. If you like DORA, use it. If you want to go with SPACE, fine. The main thing is that you don't go in expecting insight to flow to you. Those are investigation metrics and of little utility.
"What about AI," you're probably wondering. Well, AI is a tool and a pretty neat one but it's more like developers using IDEs instead of simple text files than it is anything else. Just as trying to extract useful information from tool usage is pointless, so is measuring AI. The exception here is token cost which I would consider typical operational hygiene. If you're wondering about the impact of AI, that is understandable, but the guidance around balancing metrics and using signals will work here as well.
I will share a list of metrics I commonly use and a brief note on how I use them.
Operational
- Cycle time - Outside of 1 standard deviation gets investigated
- Throughput - Keep within 1 standard deviation
- Change Failure Rate/Defect rate - Within 1 standard deviation
- Velocity - Within 1 standard deviation unless major change
Signal
- Sentiment in 1:1s - Could have folks quit soon
- CI goes red - Something is broken and needs to be fixed
- Deployment failure - Same
- Leader tapping a dev on the shoulder for, "Something real quick" - Poor prioritization, productivity killer
- Status update queries - Poor prioritization, priority killer, perception management issues
- Item age - Process failure, weak understanding of done, gold-plating, over scheduling
- Process exceptions, exemptions, and expedites - Process weakness, matrix org failure, poor prioritization
What Should You Report Upward?
The desire for more and better metrics or reports is constant. Most organizations favor a narrow set of things that fall into the easy-to-count category but leave a lot of questions unanswered. As you develop your competency with data you should take a few things into account.
First, your use of data will build or hurt trust. Every report is a performance. When your leadership trusts you, they take the performance at face value. When they don't, they watch it wondering what you're hiding. Use the data you have honestly and with integrity and your trust will build so much so that even if your data looks like hot garbage, your leadership will trust that you can handle it.
Second, I often advocate for what I term, "Living in two worlds." What this means is that you present the reports that are expected, but you add to it the more refined data that tells the real story that needs to be told. You'll need to do something like this because most folks just aren't good enough with data to work with a brand new set of data and they will need a lot of help understanding what you're showing. Give them what they expect and then give them what they need.
A common example that shows up for me is when I am expected to present on roadmaps. Typically they'll have the name of the project and the timelines with major milestones and risks. Oh, and they're always green or yellow. Never red. Never. What I will do is add in another section that might be key metrics or hypotheses we've validated, or new customers, or something like that. I want them to see what they expect and then engage with the new thing. That's how it starts. After a few of those my single slide may turn into two or three each layering on data that exposes deeper meaning.
If you find yourself in the unique position of having no expectations of what to present, ask yourself, "If I were in their shoes, what are the most important questions I want answers to?" You'll likely create a list like this:
- Are we on time, and if not, how early/late are we?
- What support is needed?
- Are we stuck somewhere?
- What risks are real and perceived and how are they being addressed?
- What new opportunities exist?
- What changes have taken place that I need to be aware of?
With a list like this you can now frame up the data to help assist in these answers. For example, the first question is, "Are we on time?" You might decide to use estimates for this and you would say, "We believe we are on time (Or not) based on these estimates." Now, I would then tell you to layer on actuals so you can see the difference and forecast but that is, well, advanced.
What Works Instead of More Dashboards?
It turns out I wrote a guide to help with a bit of this, but I'll distill down the main lessons. Almost every leader in every company I've worked with struggles with metrics. So you'll have to be patient with yourself and start with easy things or interesting things.
My honest advice is to begin by developing a list of signals for the decisions that are most common and important to you as a leader. They can be private and you'll grow in your ability to build out more metrics and deepen your understanding of them to create operational metrics and to know when to use investigation ones as well.
At the end of the day, if they aren't materially helping you lead your team to better outcomes, then you should jettison them.
When you feel that itch of wanting more data and more metrics, use that moment to pause and figure out what question you really want to answer. Then you can apply Signal Mapping to it, or attempt to create balanced metrics to target it.
If that is still overwhelming, I do offer a workshop to help with exactly this. Even better, it's a workshop for you and your peers so your entire leadership team grows together in their ability to use data effectively.
Frequently Asked Questions
What Are Engineering Metrics?
Engineering metrics represent the data, counts and calculations leaders use to inform decisions and monitor teams, projects, systems, and guide their own leadership. While there are many metric frameworks that exist like DORA and SPACE, the concept of leveraging these data points in leadership is what engineering metrics are about.
Which KPIs Should a Software Engineering Team Track?
A great way to start with this is in discussing this with your team. Decide what is important to both you and them, and ask, "How will we know?" Things like DORA or SPACE can help fill in blanks but are in no way complete. From there, collect the data you need to answer those questions and keep them visible so everyone can inspect and form their own opinion about the metrics and their meaning. Refine with them regularly. This will build a practice of a whole team using data to inform their decisions and efforts and that can make a huge difference.
Are DORA Metrics Enough?
No. DORA, when used well, can offer a broad-brushed picture of operational health for engineering teams or organizations. Unfortunately, many day-to-day or critical decisions that you need data and metrics for, DORA will not be able to help with. For example, DORA will never be able to tell you if you're going to be late on a critical initiative, but it may give you clues as to contributing factors.
How Many Metrics Does a Team Actually Need?
To start the practice I'd recommend starting with one metric. Don't overload its meaning though, just start with one to get a feel for how your actions show up in the data. From there work to balance that metric along leading/lagging and metrics to prevent degradation elsewhere. I also highly recommend developing a core set of signals to prevent and avoid disaster. This could be morale, outages, budgets or anything else. In general though you won't need many and certainly not a massive dashboard with filters and customization.
Why Do Engineering Metrics Get Gamed?
First, when metrics are known to be important to leaders, you can't help but want to make them look good. This is how the gaming starts. Sometimes it's a conscious effort and sometimes it isn't. Nobody wants to be singled out by their boss. So folks work in a way to make that metric look good. The problem is that it almost always means something else gets sacrificed and the intent of the metric is undermined entirely. Balance your metrics against what will be sacrificed and lagging indicators that prevent undermining.
If this kind of straight talk is useful, I write a short letter every Friday. Stories, tips, techniques, and the occasional bits of beekeeping. Join it below.