
You should not measure social engagement before you have volume, because a ranking built on a handful of posts reads noise as a finding. The honest order is a four-stage ladder: effort, goal, analysis, prescription. Secure the feed first, measure brand growth in the aggregate, and only then descend to the post level, with a minimum sample, a mature-post window, and comments weighted above likes.
No. Analysis is the reward of volume, not the starting point, and ranking posts before the feed is consistent reads the luck of the day as a finding. The work is managing content the way a sales leader manages a rep: cadence before quota, effort before result. You do not pull a rep on the number when there were no calls, and you do not open an engagement dashboard for a feed that posts twice a week. The honest sequence is a ladder of four stages, each with its own question, and together they close into a loop.
| Stage | The question it answers | What it is not |
|---|---|---|
| 1. Effort | Did the content ship, on cadence, this week? | Not whether a post performed |
| 2. Goal | Did the brand grow in the aggregate this month? | Not yesterday's like spike |
| 3. Analysis | Of what shipped, what actually earned craft? | Not a cherry-picked screenshot |
| 4. Prescription | Given what worked, what is the next post? | Not a generic "post a reel" |
Analysis is the reward of volume, not the starting point. You cannot analyze what barely exists, and a ranking of six posts measures variance, not performance.
Most teams start at stage three because it is the one with the pretty chart. They hire for content, attach a result target on day one, and build a monthly report on top of a feed that is empty three days a week. The gap underneath the report is not laziness, it is statistics: with a sample of three or six, the margin of error is wider than the difference anyone swears they saw. The post that "worked" and the post that "died" can be the same post on two different days.
Effort and goal, in that order, because you cannot measure the result of work that did not happen. Stage one measures presence, not performance: did each channel hit its cadence this week. The rule is the same one that separates a good rep from a hopeful one, cadence first, number second. A watchdog for this stage does not know how many likes a post earned, and it should not; it knows whether the post existed. Green if the channel hit its target, yellow if it is on track, red if the feed stalled.
Effort targets belong in one place, versioned, read from there, not re-agreed every Monday. And they should be anchored in real history, not aspiration. A target you never hit is not ambition, it is a lie you tell yourself every week; a view ceiling the account has never once reached does not become realistic because it is written down. Set the goal slightly above the real recent pace, not at a number that has never happened. Honest measurement starts here, by refusing to inflate the ruler.
From the first of the month, not on a rolling 30-day window, because the two look identical and are not. A rolling 30-day window, read on the 10th, glues 20 days of last month onto 10 days of this one. You see the growth number, you credit it to this month, and half of it is inheritance you dragged across the boundary. Measuring from day one resets the counter when the month turns, so what moved since then is what actually moved, with no inheritance.
Growth is the honest goal because it does not rise by luck. A single post can spike on a good day or because it was pure promotion; the follower base, the reach, the interaction on everything already published, those move because the brand is actually growing. So the goal is aggregate on purpose: did the brand grow this month, yes or no, with no cherry-picking the one good screenshot. One subtlety catches teams out: interaction accrues over time, so a post from yesterday has not yet collected the comments and saves it will gather over the week. Drop it into the average and it enters at half its value, dragging the number down, and the indicator screams that interaction fell when the posts simply have not matured. The fix is the same one that makes the next stage trustworthy: do not count newborn posts.
Enough that the sample can survive the luck of the day, which in practice means a floor on count and a window that excludes immature posts. Two rules carry this, and both live in the data, not in anyone's judgment:
Why this matters is visible in any small ranking. Suppose the best post of the week drew 22 interactions, the second drew 13, and everything else sat near zero. Crowning a winner there is reading noise as a discovery, because the top post and the dead one may have each caught a single lucky or unlucky day. The analysis machine can be ready and run clean; what is not ready is the volume that turns a ranking into a conclusion instead of a well-packaged guess. That is exactly why the ladder starts at effort and not here.
Because a like is the lazy metric and a comment is evidence of craft. A like is the automatic thumb of someone who already follows you, almost a reflex, and it inflates hardest on a post that is pure promotion. It tells you a finger moved, not that the content was good. A comment, a repost, a save: those cost the other person time and exposure, so they are the signal that the work made someone stop and spend energy. An honest ranking weights that difference instead of summing everything into one vanity number.
In the ranking that decides what worked, a comment is weighted 3x a like. It is a deliberate, debatable choice, but it is explicit and it lives in the query, not in anyone's mood.
The point of putting the weight in the data is that it is defensible and repeatable. Anyone can argue the multiplier should be two or four; what they cannot do is quietly change it per report. The ranking means the same thing every week because the rule that produces it is written once, the same principle behind reading the UTM instead of an inferred source in why your CRM under-reports paid leads.
SQL calculates the number; the model only writes the sentence. This is the boundary that recurs in everything built with an LLM, and it is not optional. The query counts the posts, computes the growth, ranks the formats; the model receives the finished number and does the language work, choosing the words, picking the status color, making it readable. The model never estimates, never "thinks" four posts shipped, never does arithmetic. If you let a model do the math, one day it does the math wrong inside a beautiful, convincing sentence, and a decision gets made on an invented number.
This is the social-metrics case of a general rule: the business logic belongs in the data, not in the prompt. A metric computed in the model is a metric that drifts between runs, the same failure described in why AI agents give different answers on the same data. Keeping the calculation deterministic and the writing generative is one more instance of attaching meaning to the data itself, covered in context engineering for AI agents.
Nekt is a data platform that gives AI agents governed context, and a social scoreboard is a clean example of the three parts it covers.
The value is not the agent that paints the status dot. It is the layer that guarantees the dot tells the truth, because without it the scoreboard lies with confidence.
The calendar month, measured from the first. A rolling 30-day window mixes part of last month into the current count, so a mid-month reading inherits growth that did not happen this month. Resetting on the first keeps the number honest: only what moved since day one counts.
Enough to survive the luck of the day. A practical floor is at least 15 posts for a format, and counting only posts mature enough to have collected their interaction, roughly those published between 37 and 7 days ago. Below that threshold, a ranking measures variance, not performance, so the honest move is to withhold the conclusion rather than dress a guess as a finding.
No. The calculation should be deterministic, done in SQL, and the model should only write the text around the finished number. A model that does arithmetic will eventually be wrong in a fluent, convincing way, and decisions made on an invented number are worse than no number at all.