The Likert scale is the most widely used form of rating scale in empirical social research — introduced in 1932 by the American psychologist Rensis Likert, and since then developed further in countless variants. It measures unobservable constructs: attitudes, satisfaction, importance judgments, purchase intentions. What methodology demands is simplicity for respondents and precision for analysis. These two goals stand in tension, and the following five decisions are what balances them out.
Aspect 1 — how many response categories?
Textbooks and method collections offer no canonical number, but rather a range: four to nine categories. The most common designs in practice are the five-step and the seven-step scale. Which you choose depends on three factors.
First, the differentiability of the language. A scale has to be namable. With five steps you can easily write "disagree — somewhat disagree — neither nor — somewhat agree — agree." With seven steps it gets tight; with nine steps it gets artificial. What you can't clearly distinguish in language, you can't clearly distinguish cognitively either — and then respondents guess.
Second, the medium. On a smartphone, long horizontal scales either become too narrow or require scrolling. In a telephone survey, all categories have to be read out — seven is then the pragmatic maximum that respondents can still remember. In a paper-and-pencil survey, space considerations come into play as well.
Third, the analytical purpose. The finer the scale, the higher the statistical resolution that can theoretically be achieved. But: each additional step yields only diminishing marginal returns, because respondents simply can't differentiate as finely as the tool can. Correlation-based analyses benefit from five to seven steps; going further rarely brings real analytical gain.
Aspect 2 — balanced or not?
A balanced Likert scale contains equally many agreeing and disagreeing categories. An unbalanced scale shifts the weight — for example three positive categories ("strongly agree" / "agree" / "somewhat agree") against just one negative ("disagree"). At first glance the difference looks harmless; in the data it isn't.
Unbalanced scales create an invisible pull toward the overrepresented side. Respondents take their cue from the choices offered and read an expected answer pattern out of what the scale presents — "the positive answer is apparently the normal one." This bias is well documented in the methodological literature and can markedly distort a study's findings.
Balance also means that the semantic distances between the categories are evenly spread. "Strongly agree — mostly agree — neither nor — disagree — strongly disagree" looks balanced (two positive, two negative, one middle), but on closer inspection it isn't: between "neither nor" and "disagree" a step is missing that is present on the positive side. Cleaner would be "strongly agree — agree — neither nor — disagree — strongly disagree" — the same step between all categories.
When in doubt: a collegial read-through before the field phase. Imbalances in scales, once you know what to look for, are found in thirty seconds. In the dataset later, no longer.
Aspect 3 — an even or odd number of steps?
An odd number of categories has a neutral middle ("neither nor"). An even number has none — respondents have to commit to one side. The choice is a deliberate design statement.
A neutral middle makes sense when you assume that a substantial minority of respondents genuinely holds no position — on abstract topics, on questions about barely known objects, on sensitive or politically charged content. Drop the middle and you risk a "tendency to the middle" as a forced-choice artifact: respondents then pick the next category up or down without any position-based anchor.
An even number of steps without a middle makes sense when the construct requires taking a stance. For importance ratings ("important — somewhat important — somewhat unimportant — unimportant") you want to know which way the tendency leans. A middle position would usually not be informative here in terms of content.
A good rule of thumb: if you can't decide, ask yourself — would I be able to use this data if forty percent of the answers landed in the middle? If yes, leave the middle in. If not, use an even number.
Aspect 4 — mandatory answer or "don't know"?
Related to the question of the middle, but to be distinguished from it, is the question of an explicit "don't know" option. It captures a different cognitive state: not "I don't care" (that would be the neutral middle), but "I have no information or opinion on this point."
The temptation to leave out "don't know" is understandable — you want answers, not refusals. But: forced answers from respondents who don't know the topic are noise. They don't improve the analysis; they only make it more prone to interference.
A pragmatic rule of thumb: use "don't know" when the question presupposes specialist knowledge or touches on content on which some respondents plausibly have no opinion (politics, scientific topics, personally sensitive areas). Do without it when the topic is universally accessible (satisfaction with one's own workplace, one's own consumption preferences).
A second variant: "no answer" as an explicit refusal option. Methodologically this is narrower than "don't know" — respondents know the answer but don't want to disclose it. In QUESTIONSTAR, both options are analyzed separately, so you can distinguish between them afterward.
Aspect 5 — label all steps or only the endpoints?
Three variants occur in practice: fully verbally labeled scales, scales labeled only at the endpoints (endpoint scaling), and pure numeric scales without words. The methodological research on this question is clear-cut — and in a direction that surprises many at first: for respondents' answering behavior it makes practically no difference whether only the two outer steps or all scale points are labeled.
The choice is therefore less a methodological one than a pragmatic one. Full labeling gives every step a linguistic anchor — that helps respondents who read quickly, and makes the questionnaire easier for outsiders to interpret. The price: you have to find a plausible label for each step, and from five to seven steps on, the language inevitably becomes artificial.
Endpoint scaling ("fully applies" on one side, "does not apply at all" on the other, with empty points in between) requires respondents to supply their own internal yardstick. This works well with culturally established scales (school grades, NPS) and is often the cleaner solution in multilingual studies — the semantic translation of middle steps like "applies somewhat less" is rarely exactly equivalent across languages, whereas the poles are.
Pure numeric scales without labels are a special variant. They work when the scale is culturally established (NPS from 0 to 10, school grades from 1 to 6) and needs no further explanation. Outside these established cases they lead to uncertainty about what the individual numbers are supposed to mean.
Conclusion
Likert scales aren't a window you simply open. Each of the five decisions above has consequences — sometimes small, sometimes large. Think them through in advance and you get data you can work with. Skip them and you collect answers whose informative value is decided only at the analysis stage, and then often for the worse.
If you're unsure for a specific project which settings suit you, write to us. We answer methodological questions even when no contract is on the table.
References
- Baur, Nina, and Jörg Blasius (eds.): Handbuch Methoden der empirischen Sozialforschung. Wiesbaden: Springer Fachmedien, 2014.
- Jacob, Rüdiger, Andreas Heinz and Jean Philippe Décieux: Umfrage: Einführung in die Methoden der Umfrageforschung. De Gruyter Oldenbourg, 2019.
- Malhotra, Naresh K., and David F. Birks: Marketing Research: An Applied Approach. 2nd European Edition. Pearson Education, 2006.
- Schumann, Siegfried: Repräsentative Umfrage — praxisorientierte Einführung in empirische Methoden und statistische Analyseverfahren. 7th edition. De Gruyter Oldenbourg, 2019.
- Thielsch, Meinald T. (ed.): Praxis der Wirtschaftspsychologie. Themen und Fallbeispiele für Studium und Anwendung. Volume 2. MV Wissenschaft, 2012.
- Weinreich, Uwe, and Eike von Lindern: Praxisbuch Kundenbefragungen — repräsentative Stichproben auswählen, relevante Fragen stellen, Ergebnisse richtig interpretieren. mi-Fachverlag, 2008.
