Views Measure Views

Nine years in, I finally worked out what else to count.

A writer I follow, Sylwia Laskowska, recently published a post about accidentally becoming a blogger. She has been writing on DEV for about a year, and the numbers attached to that year are impressive: hundreds of thousands of views and tens of thousands of followers.

I have been on DEV for more than nine years. This morning I am at 48,408 total views.

Before I go any further, I want to be clear about something, because the essay that usually follows a comparison like that is insufferable. Building a large readership for accessible, useful developer writing is genuinely difficult, and doing it in a year is more difficult still. I would like more people to read my work. I can learn a great deal from writers who are better than I am at audience building, topic selection, accessibility, and community participation.

So this is not a piece about why small numbers are secretly good. It is a piece about what I spent nine years failing to separate.

The number that stopped me using one of my metrics

I recently pulled nine years of my own data out of the DEV API, mostly to make a chart of follower growth with my publication dates marked on it, so I could see which posts moved the line.

The chart was useless, and the reason is instructive.

In 2026 I gained roughly eighteen thousand followers. My 2026 posts have about fourteen thousand views between them. You cannot acquire eighteen thousand readers from fourteen thousand page loads. The daily follow rate sits around 130 and does not respond to whether I publish anything, and about 37% of the usernames carry auto-generated hex or numeric tails. It is reciprocal-follow farming, it is endemic, and it has nothing to do with me or my writing.

Here is the cleanest version of it. Since the middle of September, my follower count has gone from 18,148 to 21,048. Over the same fifteen days, my total view count went from 46,774 to 48,408.

Two thousand nine hundred new followers. One thousand six hundred and thirty-four new views.

I gained nearly twice as many followers as readers, and a follower is supposed to be a reader who liked something enough to want more. That number had been sitting on my profile for months looking like evidence of something.

A page view is at least an honest measurement. Somebody loaded the page. That is real information, and I think “vanity metric” is an unfair label if it is taken to mean meaningless.

The trouble starts when we quietly change the claim from “this post received more views” to “this post was more successful.”

Successful at what?

A page view is real. It just is not the whole story.

In corporate content the answer to that question is usually explicit. A post might exist to attract someone searching for a problem, introduce a product, move that reader toward a trial, and eventually help create a customer. Ten thousand views with no downstream behaviour may be worth less to that company than five hundred views that put twenty qualified developers into the funnel.

Personal writing has a funnel too, just a much vaguer one. Someone reads one article, encounters another a month later, starts recognising the name, follows, leaves a substantive comment, references the work elsewhere. The page view was real. It was not necessarily the outcome.

It also helps to remember that a broadly useful JavaScript tutorial and an article about authority boundaries in AI-generated software are not competing for the same reader. One is relevant to a large share of a developer community. The other starts with a much smaller pool of people who already care about capability security and the distinction between correctness and authority.

That is not so different from comparing the audience for a popular fantasy novel with the audience for a presidential memoir. Both books can be excellent. Both can do exactly what their authors intended. Their potential readerships are still radically different.

Audience size is partly a property of the artifact and partly a property of the market around it. That sounds obvious about books. Writers forget it remarkably fast while staring at a dashboard.

An article can also have a desired consequence without a button under it. Sometimes I want someone to challenge the argument. Sometimes I want a developer to recognise a problem in their own system. Sometimes the entire intended outcome is that a reader leaves with a question they were not asking ten minutes earlier. That still gives me something to evaluate against. It just refuses to appear as conversion=true in an analytics dashboard.

Quality, audience fit, and distribution are different problems

I have come to think of online writing as at least three separate problems.

Writing quality is the craft problem. Is the argument coherent? Is the explanation useful? Did I do the research? Is there something here worth another person’s time?

Audience fit is the relevance problem. How many people where I publish are likely to care about this subject? How much prerequisite knowledge does it demand? Can someone scrolling a feed see immediately why the question matters to them?

Distribution is the discovery problem. How does the article reach those people? Search, followers, newsletters, speaking, community participation, platform curation, links from other writers, or some combination.

Those interact, but they are not interchangeable. An article has to clear all three, and failing any one of them produces the same disappointing number for completely different reasons.

Flowchart. An article passes through three decision gates in sequence: quality, then audience fit, then distribution. Failing the quality gate leads to "nobody finishes it." Failing audience fit leads to "few people care." Failing distribution leads to "nobody sees it." Clearing all three reaches readers.

That is the diagnostic value of separating them. Three posts can land at 80 views apiece and need three entirely different responses. A good article can have poor audience fit. An accessible article can have excellent fit and no distribution. A technically modest piece can answer a question a hundred thousand people are asking today. A strong argument can address a question five hundred people know they have.

For most of my writing life I concentrated almost entirely on the first variable and assumed distribution would sort itself out. Sometimes it did. Frequently it did not. My own numbers say that plainly: about 4% of my traffic comes from search, which for someone whose best-performing historical work is evergreen reference material is a distribution problem rather than a quality one.

I did not start with a content strategy

Nine years ago my public technical writing grew out of databases, because databases were the work I was doing and the community I was in. I did not sit down with a personal-brand document. I wrote about what I knew, what I was learning, and what developers were asking about.

Those articles accumulated into something larger. People began associating my name with certain subjects, questions led to more articles, and some of that work became durable enough to keep attracting readers years later. My database writing eventually included MongoDB’s Building with Patterns, one of the most heavily visited bodies of content I worked on there.

That history matters when I wander. If I publish one Rust article today, it does not arrive with nine years of association between my name and Rust behind it. That does not mean developers are uninterested in Rust, or that my readers dislike it. It may simply mean I have not given a Rust audience any reason to know who I am.

Those explanations imply completely different responses. If I wanted Rust to become a real part of my writing, one underperforming article would tell me almost nothing. I would need several useful pieces, participation in that community, and time. If I do not want that, the article stays an interesting experiment.

“This article performed poorly” is an observation. “There is no audience for me here” is an interpretation.

Personal writing gets to discover its strategy

I have spent enough of my career around corporate developer content to know that content strategy matters. A company generally knows why it is publishing: which developers it wants to reach, which capabilities it needs explained, which search terms it wants to own. The strategy should exist before anyone fills the editorial calendar.

Personal writing is stranger. You can chase a question because it bothered you on Tuesday. You can abandon a series when you have nothing else useful to say. You can spend weeks on a technical experiment and then publish something ridiculous because you started wondering whether all the photographs on your phone technically make it heavier.

You can also discover the strategy after you have written enough to see the pattern. That has increasingly been my experience. Articles I thought were about AI verification, provenance, missing information, audit evidence, memory, and authority turned out to be different views of the same few questions. I did not design that body of work and then manufacture articles to fill it. The writing is how I found it.

The conversation that prompted this piece made me realise that is less different from corporate strategy than I assumed. Sylwia described writing mostly by intuition while still making choices about what she wants to be known for, which audiences interest her, which adjacent topics fit, and which opportunities she ignores.

That is a content strategy. It just has a governance structure of one.

A company may need content, DevRel, product marketing, SEO, and leadership involved in deciding whether a newly discovered audience matters. A personal writer can notice something in the comments on Tuesday and run the experiment on Thursday. The strategic question is nearly identical. The path from observation to decision is not.

Signals are not instructions

This matters because audiences talk back. A recurring question in the comments may reveal an adjacent audience you did not know you had. Search traffic may show people finding an article for a reason you never anticipated. A series may attract platform engineers when you thought you were writing for application developers.

That is useful information. It is not an order.

A writer can discover that beginner tutorials have an enormous reachable audience and still decide not to build a body of work around them. A company can discover that a group of users loves a product for an unexpected use case and still decide that market does not fit the strategy.

Discovering an audience is not the same as deciding to serve it.

This is where “write more of whatever did best last week” collapses. It is the same loop with one step deleted.

Cycle diagram. Writing leads to publishing, which produces signals: views, comments, search terms, and who shows up. The signals reach a decision point asking whether this is where the writer wants to go. Yes leads to building an audience there. No leads to noting it and leaving it alone. Both paths feed into strategy, which returns to writing.

The diamond is the part that gets skipped. Without it the loop still runs, it just runs on autopilot, and the writer ends up somewhere chosen by whatever the feed rewarded in a given week.

Metrics tell you what happened. Readers reveal opportunities you did not know to look for. Neither one gets to decide what you want the work to become. The feedback loop needs interpretation.

I eventually wrote down some rules

Once I could see a body of work forming, it became tempting to turn every passing thought into another strategic article. So I wrote myself a gate.

For the deliberate part of my technical writing, I now ask whether there is a disputable claim, whether investigating it will put pressure on an actual artifact, whether I have standing through a project or experiment, whether it advances rather than repeats the larger body of work, and whether a reader should do or question something differently afterward. I also want to know what would falsify the claim before I start assembling evidence for it.

The artifact-pressure test is deliberately hard. If investigating an idea will not change code, a specification, an ADR, a schema, or a demo, it probably does not belong in that stream. The investigation also has to be capable of failing. Building fixtures that encode what I already believe proves very little.

That produces a shape I have become fond of:

Here is what I thought.
Here is what would have convinced me I was wrong.
Here is what I built.
Here is what happened.
Here is what changed.

Those are not my rules for everything. I deliberately keep another lane with almost no gate at all: career observations, language experiments, satire, community responses, project archaeology, and pure curiosity only need to be worth writing.

A publishing strategy should help me recognise strong work. It should not make me ask permission before being curious.

Some of it is luck

There is one variable I cannot put into a strategy with any confidence.

I can study years of data and conclude that Tuesday at 9:00 a.m. Pacific is the right time to publish. That says nothing about whether it is right for this article. A major news event may take the morning. Three other posts aimed at the same readers may appear within the hour. A moderator may promote something. A Gem may land. Another writer with a large audience may link to you. None of that says anything new about the quality of the work, and each can transform its distribution.

Strategy does not eliminate luck. It changes the conditions under which luck operates. You can improve the writing, understand the audience, participate in the community, publish consistently enough that people know you exist, and make the work easy to find.

You can engineer more opportunities for an outcome without engineering the outcome itself.

And when luck does hand you an unexpected success, strategy comes back. A surprising audience appearing is another signal, not a mandate. You still have to decide whether it points somewhere you want to go.

What I actually track now

I still look at views. Pretending otherwise would be silly. If one article gets 5,000 and another gets 40, I want to know why. I just no longer think the first was 125 times more successful.

The outcomes I care most about do not fit in a platform dashboard. A reader challenges a claim and I change the model. A comment exposes a missing invariant. An experiment breaks the answer I expected to publish. An article changes code, a specification, or an architecture decision.

That has become much less theoretical lately.

I published an article that began with a refractometer and a batch of homemade wine, arguing about the difference between a measurement and the state we infer from it. The comments pushed it considerably further: version the correction rule, spend more measurement budget on consequential baselines, decide how conflicting instruments get adjudicated before seeing the readings, distinguish evidence that survives a restart from evidence that dies with the process.

An article about authority boundaries in AI-generated code did the same thing. Readers pushed on capability lifetime, consumable authority, semantic authority diffs, denied-call telemetry, and who is permitted to modify the authority boundary itself.

Those comments did not just increase engagement. They changed the model. Those are propagation effects rather than distribution metrics, and I have started tracking them separately: research-induced change, engagement from people with standing outside my field, independent use of a concept, substantive challenges and extensions, and whether writing has started consuming so much time that the projects supplying it have stopped moving.

My numbers support the split more cleanly than I expected. Across nine years I have 1,372 reactions and 683 comments. Roughly one comment for every two reactions is a strange ratio, and it is concentrated almost entirely in recent work. My 2017 tutorials pulled more than twice the traffic of everything I wrote in 2026 and produced 26 comments in three years. The 2026 essays, with a third of the traffic, have produced hundreds. One of them has a comment thread 35 replies deep.

One body of work got found and skimmed. The other gets read and argued with. For years I evaluated both with the same number.

A related consequence: a technically unsuccessful investigation can make a successful article. If I start with a hypothesis, state what would falsify it, build something capable of producing an answer I do not control, and find out my model was wrong, that is useful. Possibly more useful. It tells the reader something, it changes the artifact, and it often exposes a better question than the one I started with.

I have a line in my current content plan that says missing a publishing day beats manufacturing a weak post. Nine years ago I am not sure I would have been comfortable with that.

Nine years later

I am still working this out. I still publish things and wonder whether anyone will care. I still occasionally write something I expect to perform well and watch it vanish. I still publish something on a whim and find it landed exactly where it needed to.

The difference is that I now have more ways to recognise success when it shows up wearing something other than a large number.

A successful post might reach 30,000 people. It might produce a conversation that changes the next article. It might expose a flaw in an architecture. It might give someone language for a problem they were already having. It might lead to code. It might turn out to be part of a body of work whose shape I could not see when I wrote the first piece.

And sometimes it might simply be an article I wanted to write, written well enough that I am still happy to have my name on it years later.

Views measure views. They are a real measurement of a real thing, and they do not independently measure rigour, usefulness, influence, audience fit, changed behaviour, changed artifacts, or whether the writing moved me toward a better question. Sometimes those correlate. Sometimes they do not. Mine told me almost nothing for nine years, and the twenty-one thousand followers told me less.

The useful question is not “Was this post successful?”

It is “What did I want this post to do, and what happened because I published it?”

Those are much harder numbers to put on a dashboard. I think they are also the ones worth learning to notice.


Thanks to Sylwia Laskowska for the conversation that prompted this, and for the encouragement to write some of it down.

Facebooktwitterredditlinkedinmail

The Most Useful Line on Your AI Cost Report Is the One You Can’t Explain

Attribution, allocation, and why “unknown” belongs in the schema.

This piece grew out of a comment thread on Sarvar Nadaf’s Per-Agent Cost Tracking for Multi-Agent AI on AWS. The schema below was worked out in that conversation, in public, and it is better for it. Where a specific idea came from the exchange, I have tried to say so.


Most AI cost dashboards answer one question well: how much did this run cost. Total tokens, model spend, per-agent spend, latency, tool usage. Those are real and useful numbers, and for a while they are enough.

They stop being enough the moment your system becomes a composition. Once a request flows through a retriever, a knowledge graph, three specialist agents, and a supervisor that synthesizes their output, the total tells you almost nothing about what to change. A run can be correct, return HTTP 200, look healthy in every latency-and-errors panel, and still cost forty percent more than an identical run that produced the same answer. The overspend is real. It is just not anywhere you are looking.

To find it, you have to stop asking where the money was spent and start asking what caused it to be spent. Those are different questions, and the gap between them is the whole subject of this piece.

Where a Cost Is Incurred Is Not What Caused It

Consider a retrieval operation that costs $0.004: searching, ranking, fetching. That is the direct cost, and it is easy to attribute. It happened on that span, you can measure it, done.

Now suppose that retrieval returned 20,000 tokens, and all of them were hydrated into a supervisor’s context on the next step. The supervisor then costs $0.009. How much did the retrieval really cost?

The direct answer is still $0.004. But that is no longer the interesting answer, because the retrieval also caused cost somewhere else. It inflated the supervisor’s context, and some portion of that $0.009 exists only because the retrieval handed it too much material. The cost was incurred at the supervisor. It was caused, in part, at the retriever.

This gives you two distinct dimensions, and a useful cost model has to carry both:

  • Where the cost was incurred. This is just the span. Directly observed, low ambiguity.
  • What caused or contributed to it. This is the interesting axis, and it is the one no aggregate dashboard shows.

A record that captures both might look like this:

span_id: retrieval-104
direct_cost: $0.0040
primary: retrieval
returned_tokens: 20000

downstream:
  span_id: supervisor-105
  attributed_cost: $0.0021
  caused_by: retrieval-104
  attribution_method: proportional

The dollars are single-counted. We do not charge the $0.0021 twice. The supervisor genuinely incurred it; the retrieval genuinely contributed to causing it; and the record says both without inventing money. What we have added is lineage: a link from a downstream cost back to the decision that helped produce it.

The Hard Part Is Honesty About How You Know

Here is the question that breaks naive versions of this: how do we know retrieval-104 actually caused $0.0021 of the supervisor’s cost, and not some other amount?

Sometimes you can measure it. If you have a controlled comparison where the only meaningful change is that retrieval result, the delta is real evidence. Supervisor costs $0.006 without the retrieved material and $0.009 with it, so roughly $0.003 of downstream cost is attributable to that retrieval. That is measured causation, and it is the strongest claim you can make.

Most production traces do not give you that. In a real run the supervisor is carrying system instructions, conversation state, the outputs of other agents, tool results, and the retrieved material, all at once. There is no clean counterfactual. So you fall back on allocation: split the supervisor’s context cost proportionally by the tokens each source contributed. That is a reasonable method. It is not measurement, and the receipt must not pretend it is.

This is why the single most important field in the whole schema is not a dollar amount. It is this:

attribution_method:
  - measured_delta
  - proportional
  - estimated
  - unknown

That field is what keeps the entire model honest. It stops a proportional guess from masquerading as measured causation. With it, a line can say “retrieval span 104 contributed an estimated $0.0021 of downstream context cost, allocated proportionally by hydrated token share,” and every word in that sentence is defensible, because the method is stated. Without it, the same $0.0021 acquires a precision the evidence never earned.

Resist collapsing this into a confidence score. A number like confidence: 0.82 feels rigorous and gives you nothing, because now you have a second number whose provenance you have to go investigate. measured_delta, proportional, estimated, and unknown each tell you why you are entitled to believe the figure. The method is the provenance. A score would hide it.

Show the Method Where the Decision Is Made

A natural instinct is to keep the attribution method as drill-down metadata, out of the main view, so the report stays clean. That instinct is wrong, and it is wrong for the same reason aggregate dashboards are wrong: it makes two different claims look equivalent.

The method belongs inline, next to any attributed cost, with one sensible exception. A directly observed cost carries no ambiguity and needs no method tag:

retrieval-104   RETRIEVAL   $0.0040

There is nothing to disclose there; it was measured on the span. But the moment a number is attributed rather than observed, the method has to ride along:

retrieval-104 -> downstream CONTEXT   $0.0021   proportional

Drop the word proportional and that $0.0021 visually becomes as solid as the $0.0040 above it, which is a lie of formatting. A report that hides the distinction between what it measured and what it allocated has committed the same sin as the dashboard that only shows a total. If two numbers make materially different claims, the interface must not make them look the same.

So the main report shows amount, category, direct versus downstream, and method. The drill-down holds the evidence behind the method: hydrated token counts, comparison runs, parent-child span references, the assumptions the allocation rests on. The decision surface stays readable; the receipt is one click away, not dumped into the table.

“Unknown” Is Not a Gap in the Accounting

Every honest version of this schema has to allow unknown as a real value, not a placeholder you feel bad about. And once you sit with it, the unknown rows turn out to be the most useful rows in the report.

A high downstream cost with unknown attribution is not incomplete bookkeeping. It is the system telling you exactly where your observability boundary stops letting you explain its own behavior. It is pointing at the place where you cannot yet answer “what caused this,” which is the place most worth instrumenting next. A tidy report with no unknown rows has usually not achieved understanding. It has hidden its ignorance behind confident allocation.

There is also a real reason unknown is sometimes the only honest answer: context is not additive. An extra 5,000 tokens of context does not simply add a proportional slice of cost. It can change caching behavior, alter the reasoning path the model takes, or change how much output the model generates downstream. When that happens, token share and cost share stop mapping to each other cleanly, and any proportional number you report is a polite fiction. In those cases the schema should say unknown and mean it, rather than allocate a figure it cannot defend.

Read that way, the report stops being a statement of where you paid and becomes a map of two things at once: what you can explain about your spending, and where your ability to explain it runs out. The second map is the one that tells you what to build.

What This Actually Costs to Build

The reason this is not a research project is that most of the structure already exists. If your traces are parent-child spans, the causal lineage is physically present already; a downstream cost sits under the decision that produced it. You are not inventing a new tracing mechanism. You are making the attribution semantics explicit on top of a trace you already record.

Concretely, that is a small number of additions. Stamp a primary category on each span. Add a contributes_to link and a cause on spans that produce downstream effects. Add the attribution_method on any attributed cost. Then roll the report up along two axes, category and direct-versus-downstream, and let unknown be a first-class row rather than a swept-under one. The intelligence is not in the plumbing. It is in having the honest method field and being willing to publish the unknown rows.

One warning from the same conversation that produced all this: how you record and how you attribute are coupled. Change the way spans are emitted and you can silently break the logic that reads them. The defense is the same one that makes the whole model trustworthy, a known-good baseline you compare against, so that when your instrumentation shifts under you, the numbers move and you notice.

The Point

We have spent a lot of effort making AI spend visible. Better token counts, per-agent breakdowns, nested traces. All of it answers “how much.” Almost none of it answers “why,” and “why” is the only version of the question you can act on.

The move from one to the other is not a bigger dashboard. It is a small, honest schema: separate where a cost was incurred from what caused it, state the method behind every attributed number, and treat the costs you cannot explain as signal rather than embarrassment. Do that, and the report stops telling you what you spent and starts telling you what to fix, including, in the unknown rows, where to look first.

Even the cost report, it turns out, needs provenance. Not just the amount and the category, but how sure you are about who to blame. That last column may be the most useful one on the page.


With thanks to Sarvar Nadaf, whose post and the conversation under it produced this schema, and to the commenters in that thread who pushed on the baseline and the propagation. The receipt is better for the argument.

Facebooktwitterredditlinkedinmail