River
Y CombinatorBacked by Y Combinator

People & Exec SupportFree

Open-Ended Survey Comment Theme Analysis

Comment analysis ranks themes by mention count, so a theme 34 people repeated in every box outranks the one 58 people raised once.

Start here

River counts people, not mentions. In the worked example 611 free-text fields from 178 respondents hold 1,187 theme mentions across 31 themes, because a comment averages 2.27 topics. The loudest theme has 118 mentions and 34 people behind it, who raised it in nearly every box they filled. The theme ranked fifth on mentions has 61 mentions and 58 separate people. Re-rank on people and the first drops to seventh while the fifth rises to second.

Then every theme is classed by whether anyone here can change it. Fourteen of the 31 are actionable and get a named owner. Nine are structural, following from the funding stage, the geography, or a decision already taken. Eight are neither: they ask for something that already exists and was never announced, and those eight reach 96 of the 178 respondents. That bucket costs an email to close and it is routinely the largest.

Written for whoever has to present this round on Tuesday and is holding 611 rows with no analyst and no week to spend on them. What comes back is one sheet of themes and one document that says which of them anybody here can change. It reads the round an engagement survey pack fielded, hands rating themes to a calibration session, and sends pay themes where they are actually decided, in the compensation review cycle. The same counting problem in hiring is interview scorecard synthesis.

A survey comment export being themed by how many people raised each topic
Written for the week between survey close and the leadership readout.

The denominator is right in the privacy layer and wrong in the analysis

Your platform already knows the correct unit. Culture Amp documents its comments group minimum as the number of people in a group who submit a response, noting that it is not the number of comments written that is the minimum. People is the denominator that protects the respondent. Then the same export gets themed by mention volume, which counts the same person up to four times, once per open question. A theme that eleven determined people mention in every box outranks one that fifty people raised once, and the readout inverts.

The second correction is for who did not answer. Non-response is not evenly spread: at Marloe Logistics the teams whose last commitment was dropped are 23.3% of headcount and 43.5% of the silence. OPM applies exactly this correction to the federal employee survey, adjusting weights for non-response within subgroups because biases can occur when some subgroups participate more or less than other subgroups. Re-weighted, one theme moves from 21.9% of respondents to 27.3% of the population.

A handful of comments must leave the pipeline before anything is counted. Six of the 522 here describe conduct rather than sentiment, four of them naming a person. Under the EEOC's sex discrimination guidelines an employer is responsible for harassment between fellow employees where it knows or should have known of the conduct, absent immediate corrective action. A comment box is a way of knowing. Aggregating those six into a theme called communication is how a company loses them for a quarter.

How it works

  1. Send the export

    The comment export as the platform produced it, every open question and every demographic column.

  2. Send the population

    Headcount by team, level, location and tenure, responses by the same cuts, and your reporting minimum.

  3. What happened last time

    What was promised after the last round, what landed, what died, and which decisions are closed.

  4. Get the ranking twice

    By people and by weighted population, split three ways, with an owner on everything actionable.

What you get

  • Every theme ranked by distinct people, with the mention count shown beside it for contrast
  • A second ranking weighted for who did not answer, and the teams that moves
  • Each theme classed as actionable, as structural, or as already decided and never announced
  • Sentiment per theme per segment, with bimodal themes flagged rather than averaged flat
  • Verbatims only where the segment clears your comments minimum, paraphrase everywhere else
  • Comments describing conduct routed out on day one, never counted into a theme

Common questions

Why not just rank themes by how often they come up?

Because a comment export has one row per question, not one per person, so a determined respondent counts up to four times. The top theme here has 118 mentions and 34 people. The one ranked fifth has 61 mentions and 58 people. Rank on mentions and you spend the quarter on the smaller group.

What does splitting actionable from structural actually change?

It stops the readout promising things nobody can deliver. Nine of these 31 themes follow from the funding stage, the geography or a decision already taken, and naming them as such is more credible than a working group that quietly dies. Fourteen are actionable and each gets an owner. The remaining eight are the interesting ones.

What is the third category?

Themes asking for something that already exists and was never announced. Eight of 31 here, reaching 96 of 178 respondents. The promotion criteria were written and approved in November; 58 people said there were none. That is a communication failure priced as a policy problem, and it is closed by an email rather than a project.

Is it safe to quote verbatims by team?

Only above your comments minimum, and usually that is fewer cuts than expected. Of 96 populated theme-and-team cells here, 19 clear a five-person minimum. The other 77 get a paraphrase with no team label, because a quote is unique text and a quote tagged to a nine-person team is effectively a name.

Why weight the results at all? Everyone was invited.

Because a census still has non-response, and it is never evenly spread. The teams whose last commitment was dropped are 23.3% of headcount and 43.5% of the people who said nothing. Weighting each cohort back to its own headcount moved one theme from 21.9% to 27.3% and from seventh place to third.

What happens to a comment that names an individual?

It leaves the pipeline the same day and goes to whoever you nominate, before any counting. Six of the 522 substantive comments here, four naming a person. A free-text box is a way of learning about conduct, and aggregating those six into a theme called communication is how they get lost for a quarter.

Does it work on a Lattice or Officevibe export instead?

Yes. What differs is column names, how many open questions there were, and where the demographic filters sit. What does not differ is the structure: one row per question per person, comments left as raw text, and a group minimum that counts people while the theme report counts mentions.

Open-Ended Survey Comment Theme Analysis

Fill in the form and your workspace opens with the work already underway.