- Cultural characteristics
- Geographic characteristics
- Demographic characteristics
- Socioeconomic characteristics
Principles of
Survey Research
Market research, survey methodology and their practical application — from research problem to results report.
Principles of Survey Research
Market research, survey methodology and their practical application — from research problem to results report.
Principles of
Survey Research
Market research, survey methodology and their practical application — from research problem to results report.
Chapter 1 · Introductory Course
Principles of
Survey Research
Methodologically sound. Practically applicable.
Market research, survey methodology and their practical application — from research problem to results report.
Contents
Introduction
Introduction
What is Research?
All systematic endeavors and efforts to acquire new knowledge for science or industry.
The search for and collection of information and ideas in response to a specific question.
Survey
A survey is one of the most popular methods of collecting primary data, in which the researcher interacts with respondents to obtain information about people's attitudes, opinions, knowledge and behaviors.
Market Research
Practical Application of Surveys
| Discipline | Application |
|---|---|
| Sociology and Political Science | Opinion research, identifying the attitudes of population groups toward socially significant phenomena, events and facts; election research (e.g. the "Sunday question"), … |
| Psychology | Personality tests, intelligence tests, identifying individual strengths and weaknesses, psychological stability, cognitive disorders, social influences, … |
| Human Resources | Measuring employee satisfaction, loyalty, potential, personality traits, leadership qualities, productivity, professional aptitude, stress resistance, work-life balance, … |
| Marketing | Market and consumer research, measuring image perception, preferences, satisfaction and loyalty (NPS), willingness to pay; segmentation, positioning, pricing, advertising tests, website usability, … |
| Science (in general) | Studying relationships between two or more variables, factors, phenomena; developing scales and methods for scientific and practical purposes, … |
| Education | Knowledge tests (multiple-choice exams), student and teacher evaluation; large-scale educational studies (e.g. PISA), … |
| … | … and many further fields of application |
Market Research Process — the "5 D's"
- Identify information needs
- Define the research problem and questions
- Set research objectives
- Assess the value of information
- Budget
- Data sources
- Research methods
- Sampling plan
- Contact methods
- Methods of data analysis
- Collect data according to the plan
- or commission an external service provider
- Analyze data statistically and subjectively
- Derive answers and implications
- Formulate the results of the data analysis
- Prepare the research report
When should you not start market research projects?
| Case | Comment |
|---|---|
| Vague objectives | When managers cannot agree on what information they need to make a decision. Market research only helps when it investigates a concrete question. |
| Fixed stance | When the decision has already been made and the study is only meant to "rubber-stamp" a preconceived plan. |
| Too late | When results are provided too late to still influence the decision. |
| Poor timing | When a product is in its decline phase, there is little point in researching new product variations. |
| Insufficient resources | It is not worthwhile to set up a quantitative study as long as no statistically significant sample is feasible — or when the finances are not enough to implement the resulting decisions. |
| Costs outweigh benefits | The expected value of information should exceed the costs of data collection and analysis. |
| Results not actionable | When, for example, psychographic characteristics are used that do not help in making concrete decisions. |
| Information not needed | When decision-relevant information is already available. |
Introduction
Types of Market Research
Market Research by Objectives
- What reasons might lie behind the declining customer satisfaction?
- What keeps first-time buyers from buying again?
- What does the historical sales trend in the industry look like?
- What are consumers' attitudes toward our product?
- Predicting the outcomes of marketing actions
- Effect of advertising spend on sales (how much does one advertising euro yield?)
Market Research by Data Source
- Generating data that do not yet exist. These data are analyzed and may, where applicable, be published by the researcher.
- Using data collected at an earlier point in time for the intended research purpose.
Market Research by Methodology
- Involves collecting and evaluating data
- Requires large amounts of data, uses statistical methods
- Aims for representativeness of the results
- Seeks to understand consumer behavior and its causes
- Focus on individuals and small groups
- Not representative — understanding one perspective, not all
Triangulation
Triangulation — combining methods within a study on the same topic.
Robson (1998) · Visocky & Visocky (2009)
ObjectiveTruth
Literature research
Survey
Interview
But in reality everything is messier.
Survey: Measurement and Scaling
Survey: Measurement and Scaling
Measurement
Measurement — assigning numbers or other symbols to characteristics of objects according to a specific, predefined rule.
Scaling
Scaling — involves a continuum on which the measured objects are placed.




Primary Scales of Measurement
- Numbers merely serve to classify the objects
- non-continuous scale
- Numbers indicate the relative positions of the objects
- but not the magnitude of the difference between them
- Differences between objects can be compared
- zero point arbitrary
- zero point uniquely fixed
- Ratios of the scale values can be computed
Some Commonly Used Scales in Marketing
| Scale | Description | Common Examples | Examples from Marketing | Statistical Measures | |
|---|---|---|---|---|---|
| Descriptive | Inferential | ||||
| Nominal Scale | Assignment of numbers to identify and/or classify objects | Passport number, football player's number, gender | Brand, gender, occupation, type of venue | Percentages, mode | Chi-square, binomial test |
| Ordinal Scale | Numbers describe the rank order of the objects, but not the extent of the differences between them | School grades, position of runners in a marathon | Preference ranking, market position, social class | Percentiles, median | Rank correlation coefficient (Spearman's ρ), Friedman ANOVA |
| Interval Scale | Allows comparison of the differences between objects; zero point arbitrary | Temperature (Fahrenheit, Celsius) | Attitudes, opinions, purchase intention, customer satisfaction, index numbers | Range, average, standard deviation | Product-moment correlation (Pearson's r), t-tests, ANOVA, regression and factor analysis |
| Ratio Scale | Zero point is uniquely fixed; allows comparison of both the distances between measured values and their ratios | Length, weight, time, money | Age, revenue, income, costs, market share | Geometric mean, harmonic mean | Coefficient of variation |
Classification of Scaling Techniques
Comparison of Scaling Techniques
The measured value of an object results from the direct comparison with another object.
Data can only be interpreted as relative positions — ordinal ordinal level of measurement only (rank order).
Each object is judged in isolation — that is, independently of other objects.
Measurement results are usually treated as intervalscaled or metric.
Survey: Measurement and Scaling
Classification of Scaling Techniques
Pros-and-Cons of Comparative Scales
Pros
- Small differences between objects can be registered
- The same known reference points for all respondents
- Easy to understand and use
- Require fewer theoretical assumptions
- Tend to reduce halo and carryover effects
Cons
- Only ordinal or rank-order level of measurement → limited choice of statistical methods for data analysis
- Data can only be interpreted as relative positions
- Impossible to generalize beyond the set of objects rated
Comparative Scales: Paired Comparison
Paired Comparison
For each pair of two objects, respondents select the one that in their opinion best fulfills a given criterion.
Below you are presented with ten pairs of beer brands. In each pair, please select the beer you would rather buy.
Comparative Scales: Paired Comparison
Paired Comparison: Pros-and-Cons
Pros
- Direct comparison and unambiguous choice
- Good for blind tests, product comparisons and MDS
- Allows calculating the percentage of respondents who prefer one object
- Rank order can be estimated (assuming transitivity)
- Possible extensions: "no difference" option, graded comparison
Cons
- Number of comparisons grows faster than the number of objects — for n objects n(n−1)/2 comparisons
- Order effects possible (influence of presentation order)
- Preferring A over B does not mean the respondent likes A
- Not very realistic for real choice situations with multiple alternatives
- Violation of the transitivity assumption possible
Violations of transitivity in paired comparison
Same respondent, same pairing — and yet contradictory answers:
From "Apple ≻ Tomato" and "Tomato ≻ Apple" no rank order can be formed — the preferences are contradictory (intransitive).
Violations of transitivity when aggregating preferences
Apple
Tomato
Orange
Tomato
Orange
Apple
Orange
Apple
Tomato
Apple ≻
Tomato2 : 1
Tomato ≻
Orange2 : 1
Orange ≻
Apple2 : 1Apple ≻ Tomato ≻ Orange ≻ Apple. Apple is simultaneously the most and the least preferred — the group preferences are inconsistent!
Comparative Scales: Rank Order Scaling
Rank Order Scaling
Respondents put several objects into an order — based on a particular criterion.
Please arrange the soft-drink brands listed below according to your preferences. To do so, first select the brand you prefer most and assign it rank 1. Then assign rank 2 to the second-best brand. Continue rating until you have assigned a rank to all brands. The last, least preferred brand must be given rank 5.
No two brands may receive the same rank.
The preference criterion is entirely up to you. There are no right or wrong answers. Just try to be consistent.
| Brand | Rank |
|---|---|
| Pepsi-Cola | _________ |
| Coca-Cola | _________ |
| Red Bull | _________ |
| Sprite | _________ |
| 7-Up | _________ |
Rank Order Scales: Example




Rank Order Scales: Examples
Rank Order Scales: Example
Orange
Kiwi
Apple
Banana
StrawberryRank Order Scales: Pros-and-Cons
Pros
- Direct comparison
- More realistic than paired comparisons
- Number of comparisons is only (n − 1)
- Easier to understand
- Takes less time
- No non-transitive answers
- Data can be converted into paired comparisons
- Good for measuring brand and attribute preferences
Cons
- Preferring A over B does not mean the respondent likes A
- No zero point — no separation between liking and disliking
- Only ordinal data
- Violation of the transitivity assumption possible (when aggregating)
Comparative Scales: Constant Sum Scaling
Constant Sum Scaling
Respondents distribute a fixed amount (e.g. points, euros, chips, %) entirely across a set of objects — according to a particular criterion.
Listed below are five attributes of cars. Please distribute 100 points across these attributes so that the number of points you assign to an attribute reflects its relative importance to you. The more points an attribute receives, the more important it is to you. If an attribute is unimportant to you, assign it 0 points. If one attribute is twice as important as another, assign it twice as many points.
| Attribute | Points |
|---|---|
| Speed | 0 |
| Comfort | 15 |
| Transmission type (manual/automatic) | 5 |
| Fuel (petrol/diesel) | 35 |
| Price | 45 |
| Sum | 100 |
Constant Sum Scaling: Example of Analysis
| Attribute | Segment 1 | Segment 2 | Segment 3 |
|---|---|---|---|
| Speed | 0 | 17 | 53 |
| Comfort | 15 | 23 | 30 |
| Transmission (manual/automatic) | 5 | 21 | 10 |
| Fuel (petrol/diesel) | 35 | 12 | 7 |
| Price | 45 | 27 | 0 |
| Sum | 100 | 100 | 100 |
Constant Sum Scaling: Example
Constant Sum Scaling: Examples
Constant Sum Scaling: Pros-and-Cons
Pros
- Can measure small differences between objects without taking up too much time
- Metrically scaled → flexible choice of analysis methods
Cons
- Results are limited to the list of objects rated — no statements about objects outside the list
- Relatively high cognitive load on respondents, especially with long lists
- Prone to arithmetic errors (e.g. distribution of 108 or 94 points)
Comparative Scales: Q-Sort Scaling
Q-Sort Scaling
A rank-order procedure in which objects are sorted into piles (with respect to a specific attribute). Used to quickly compare a large number of objects (60–140) against each other.
The number of objects per pile is limited such that all piles together reproduce the shape of a normal distribution.
The Ministry of Health has developed 25 measures for implementation in hospitals. Rank them by their effectiveness against the spread of infection — please only one measure per box.
Survey: Measurement and Scaling
Classification of Scaling Techniques
Continuous Rating Scale
Continuous Rating Scale
Respondents rate objects by marking a corresponding position on a line that runs from one extreme to the other of a given criterion.
Perception Analyzer
During the presentation of a stimulus — e.g. a TV commercial — each participant turns a dial. This creates a continuous rating in real time, second by second — aggregated across all participants and broken down by segments.
Likert Scale
Likert Scale
Respondents indicate the extent to which they agree with the listed statements — using a 5- or 7-point scale that ranges from one extreme to the other.
Below are various statements about "Real". Please indicate how strongly you agree with these statements:
| Strongly disagree | Disagree | Neutral | Agree | Strongly agree | |
|---|---|---|---|---|---|
| Real sells high-quality goods | |||||
| Real has poor service reversed | |||||
| I enjoy shopping at Real | |||||
| Real offers a mix of different brands | |||||
| The credit policy at Real is terrible reversed | |||||
| I don't like Real's advertising reversed | |||||
| The prices at Real are fair |
Important: Statements 2, 5 and 6 are reversed in wording. Before data analysis, these scales must be recoded — a higher number should always mean a better attitude.
Likert Scale: Examples
dissatisfied
satisfied
Some Commonly Used Scales in Marketing
| Construct | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| Attitude | Very bad | Bad | Neither good nor bad | Good | Very good |
| Importance | Not at all important | Unimportant | Neutral | Important | Very important |
| Satisfaction | Very dissatisfied | Dissatisfied | Neither satisfied nor dissatisfied | Satisfied | Very satisfied |
| Purchase probability(Purchase intention) | Definitely not | Probably not | Undecided | Probably yes | Definitely yes |
| Purchase frequency | Never | Rarely | Sometimes | Often | Very often |
| Agreement | Strongly disagree | Somewhat disagree | Neither agree nor disagree | Somewhat agree | Strongly agree |
Scale points from 1 (lowest level) to 5 (highest level).
Semantic Differential
Semantic Differential
A bipolar rating scale whose extremes are described by opposing adjectives. It allows the measurement of multidimensional attitudes and their profile representation.
How do you rate the appearance of “Kaufhof”? Please mark to what extent you lean more toward one pole or the other.
Kaufhof is …
Note: The negative adjectives sometimes appear on the left, sometimes on the right. This makes it possible to check afterwards whether respondents thoughtlessly always marked the same side without reading the adjectives.
Semantic Differential: Profile
Profile representation
Measures self-assessment as well as attitudes toward people or products. Each point corresponds to the mean or median of the respective scale — connecting the points yields the profile.
Example profile of an object — seven levels between the opposing poles.
Semantic Differential Scale: Example
Profile comparison
Semantic profiles of the shampoo brands “Herbal Magic” and “Elseve” compared to the ideal shampoo from the consumers' point of view.
Stapel Scale
Stapel Scale
A unipolar rating scale with 10 categories from −5 to +5, without a neutral point (0).
It is often used as an alternative to the semantic differential when no meaningful pair of opposing adjectives can be found.
Plus number = the phrase applies, minus number = it does not apply. The larger the magnitude, the stronger.
How accurately do the following phrases describe the store “Real”? For each phrase, choose a number between +5 (fully applies) and −5 (does not apply at all).
Basic Non-Comparative Scales
| Scale | Description | Examples | Pros | Cons |
|---|---|---|---|---|
| Continuous Rating Scales | Mark on a continuous line | Reactions to TV commercials | Easy to construct | Manual (non-computer-based) analysis can be very tedious |
| Likert Scalediscrete | Degree of agreement on a scale from 1 (strongly disagree) to 5 (strongly agree) | Measuring attitudes | Easy to understand, use and construct | More time-consuming |
| Semantic Differentialdiscrete | Bipolar, seven-point rating scale with opposing adjectives at the poles | Brand, product and company image | Versatile | No indication of whether the data are interval-scaled |
| Stapel Scalediscrete | Unipolar ten-point scale from −5 to +5 without a neutral point (0) | Measuring attitudes and image | Easy to construct and to use in telephone surveys | Sometimes confusing and difficult to apply |
Constructing Itemized Rating Scales
Number of Scale Categories
More categories capture finer differences — but most respondents can only handle a few categories.
Involvement & Knowledge
morewhen respondents are interested in the rating or have deep knowledge of the object.
Nature of the Objects
morewhen fine differences are characteristic of the objects.
Mode of Data Collection
fewerin telephone interviews, where scales are harder to keep track of.
Data Analysis
fewerfor aggregation & group comparisons · morefor sophisticated, correlation-based statistics.
Balanced or Unbalanced Scales
Even or Odd Number of Scale Categories
Odd number — with a midpoint
Even number — without a midpoint
The middle option attracts many undecided respondents — and those who are reluctant to reveal their opinion.
This can bias the measures of central tendency and variance.
Do we want or need “contrast” on controversial attitudes?
Forced or Non-Forced Response?
Do respondents not want to answer — or do they simply have no opinion?
"Don't know" / "Not applicable"
Skip logic
Deliberately steer so that respondents only get questions they can actually answer — rather than forcing responses.
Rule of thumb: Questions without a "don't know" tend to yield more accurate data — but only if respondents actually have an opinion.
Labeling the Scale Points
Should every scale point be labeled — or are a few selected points enough?
All or just some?
No clear evidence that labeling all points is better than only selected ones — the research shows no substantial difference.
Too much confuses
Too many, too finely differentiated labels can confuse respondents when terms can barely be told apart (e.g. "somewhat positive" vs. "fairly positive").
Decisive: clarity
Avoid ambiguity. Clearly name both poles, and with an odd number also the midpoint — it must be clear which continuum the scale represents.
Little space
Reduced labeling is especially useful with sliders or matrix questions, where full labeling quickly becomes cluttered.
How Much to Label? — Four Variants
The same question: "How likely are you to buy Product A again?"
Peaked vs. Flat Response Distribution
How extreme the endpoints are worded shapes the form of the response distribution.
Respondents avoid the extremes — the answers cluster in the middle.
Respondents also use the endpoints — the answers spread out more evenly.
Survey: Measurement and Scaling
Latent Constructs and Multi-Item Scales
Latent construct
A phenomenon (e.g. customer satisfaction) that is not directly observable or measurable.
That doesn't mean it doesn't "exist" — only that it can be inferred from other, measurable phenomena (indicators).
Latent Constructs: Hierarchy of Measurement
Multi-Item Scales: Advantages
Advantages
- Ability to assess abstract concepts
- Different facets of the construct can be captured
- Reduction of data dimensionality by aggregating many observable phenomena into one model
Multi-Item Scales: Make or Steal
Theory · secondary data · qualitative analysis
Where do you find ready-made scales?
marketingscales.com/research
Secure Customer Index
The Secure Customer Index combines three loyalty indicators into one metric. Only someone who chooses the top level (5) on Secure Customer counts as a Secure Customer on all three dimensions — the intersection. Each segment gives the share (%) of customers.
Very satisfied
Will definitely use again
Will definitely recommend
Secure Customer
Source: D. Randall Brandt (1996), "Secure Customer Index", Maritz Research
Extended Secure Customer Index by Burke Inc.
Burke extends the Secure Customer Index by two additional loyalty dimensions (five in total) and links the loyalty index measured in Period 1 with the actual Share of Wallet in Period 2 — that is, the share of spending the customer devotes to the brand.
(0 % – 100 %)
Source: Burke Inc. · http://www.burke.com/
Survey: Measurement and Scaling
The True-Score Model
The result of a measurement is not the true value of a characteristic, but only an observation of it.
Reliability and Validity
Indicates how reliably a measurement instrument measures — i.e. how consistent the results are across repeated measurements.
No random error: XR ⟶ 0 ⟹ XO ⟶ XT + XS
The measure is Cronbach's alpha (0 ≤ α ≤ 1)
Values of α ≥ 0.7 are considered acceptable
Indicates to what extent a measurement instrument actually measures the matter it is meant to measure.
That is: to what extent measured differences correspond to actual differences between the objects (quality of the measurement).
No measurement error: XS ⟶ 0, XR ⟶ 0 ⟹ XO ⟶ XT
Relationship between Reliability and Validity
The purpose of a scale is to enable us to represent respondents with the highest accuracy and reliability. We cannot have one without the other and still trust our data.
Burke, Inc. (2000–2003)
Net Promoter Score® — a predictor of company growth?
„How likely is it that you would recommend company/brand/product X to a friend, relative or colleague?"
Source: Reichheld, Fred (2003) “One Number You Need to Grow”, Harvard Business Review
Net Promoter Score®: Warning
Although the “recommendation question” is by far the best single question for predicting consumer behavior across a range of industries — it is not the best question for all industries. That is why companies have to do their homework and empirically verify.
Questionnaire
Definition & objectives
QuestionnaireA questionnaire
Questioning Techniques and Questioning Tactics
vs. Open
Questions
- Closed: choice from predefined answer options.
+ easy to analyze, no cognitive stress
− automatic, unconsidered answers - Open: answer options not predefined.
+ unlimited possibilities, taxes the memory
− complex coding, refusal possible
vs. Indirect
Questions
- Direct: the question targets the matter of interest directly.
- Indirect: the matter is inferred through an indirect formulation — sparing sensitive or hard-to-articulate topics.
indirect: "Which beverages do you prefer with meals?"
Influence of Formulation on the Answer
Same action, different formulation — opposite answers.
Source: Noelle-Neumann & Petersen (1998), p. 192 · n = 2100, p < 0.05
What Should Be Considered When Developing a Questionnaire?
Questionnaire
Asking Questions

"Not every question deserves an answer."
Publius Syrus · Rome, 1st c. BC
Avoid Ambiguity, Confusion and Vagueness
The six W's
Formulate the question in terms of who, what, when, where, why and how. Especially important: who, what, when, where.
"Which brand of shampoo do you use?" — what is unclear?
| W | Aspect | Why unclear? |
|---|---|---|
| Who | Reference person | Unclear whether only the respondent themselves or their entire household is meant. |
| What | Reference object | Unclear how to answer if several brands are used. |
| When | Reference period | No reference period given — this morning, this week, or the whole year? |
| Where | Situation / place | At home, at the gym, on vacation, on a business trip? |
Make Answer Options Complete & Unambiguous
For clarity instead of ambiguity: consider all realistic situations and prepare suitable answer options — including "does not apply" and filter routing.
2. "Are you satisfied with your current car insurance?"
Scales and Answer Options Must Be Unambiguous
Vague frequencies
Words like "rarely," "sometimes" or "often" mean something different to every respondent. Use concrete, delimited frequency specifications.
Avoid Jargon, Slang and Abbreviations
Simple words
Use simple, everyday words — no technical terms, no jargon. Every respondent must understand the question immediately.
Avoid Double-Barreled Questions
One aspect per question
Each question should focus on just one aspect. Otherwise you don't know what the answer refers to.
2. "In your opinion, is Coca-Cola refreshing?"
Avoid Leading
No suggestion
If you already want a particular answer, there's no need to ask the question. Leading wording steers the respondent.
Avoid Implicit Assumptions
Name the consequences
The answer should not depend on tacit assumptions about the consequences. Make them explicit.
Avoid implicit alternatives
Name the alternatives
Implicit alternatives are answer options that were not explicitly named. Make the counter-option visible.
Avoid Treating Beliefs as Real Facts
Facts instead of opinions
Opinions and beliefs often represent the real facts only in a distorted way. Ask about the two facts separately.
2. "Do you wear fur clothing?"
Avoid Generalizations and Estimates
No mental arithmetic
Don't force the respondent to strain their memory and their mathematical skills. Ask for the building blocks.
2. "How many members does your household have?"
Questionnaire
Overcoming Inability to Answer



Is the respondent informed?
Respondents often answer questions even when they are not informed.
Can the respondent remember?
Recall errors
Poor recall leads to omission, telescoping and creation. Ask about typical behavior, not exact counts.
Can the respondent articulate it?
Offer aids
Anyone who cannot formulate their answer skips the question or drops out. Offer pictures, diagrams or descriptions to choose from.
Questionnaire
Overcoming Unwillingness to Answer
Reduce the effort
Reduce the effort
Minimize the effort required to answer. Instead of recalling freely, let respondents select from a ready-made list.
Clarify the context
Provide context
Some questions seem inappropriate in the wrong context. Introduce them with an explanatory statement.
Explain the purpose
Legitimize the purpose
Explain why the information is needed — otherwise it seems intrusive.
Questionnaire
Handling sensitive topics
Chapter
Three principles
Funneling and Skip Logic

Place the question branched to as close to the triggering question as possible.
Arrange branching so that respondents cannot anticipate which additional information will be asked.
Example: Flowchart of a Questionnaire

Questionnaire
Design a Convincing Introduction
Pretest! Pretest! Pretest!
Test the questionnaire on a small sample before deployment — and check every aspect:
Recap
Develop a flowchart of the required information — starting from the (market) research problem.
Once the sequence is laid out, the connections become clear.
Align the collected data with the information needs.
Set a clear objective for each area — the questions follow from it.
Go back to the flowchart and ask about each piece of information:
"Do I really need to know this — and do I know what I'll do with it?"
… rather than "Nice to know, but I don't really need it."
Sampling
Dewey Defeats Truman
1948: The Chicago Daily Tribune announces the wrong election result. President Harry Truman beats Thomas Dewey — against all polls.
Reason: a biased, inaccurate opinion poll.
Sampling
Most surveys cannot survey every person. Instead, a sample is drawn and examined — this procedure is called sampling.
If the sample is drawn incorrectly, all the data is useless.
Sampling
But not everyone selected actually answers: those who really take part are the respondents.
Sampling: Two General Methods
The sample is drawn based on the personal judgment of the researcher — often at random (convenience sample, e.g. passersby in a shopping mall).
Usually inexpensive; allows a rough estimate of the population parameters.
The sample is selected based on the principle of randomness.
Allows statistical techniques to determine the accuracy of the estimated population parameters as well as to assess their confidence intervals.
Sampling Techniques
Sampling
Convenience Sampling
In convenience sampling (selection at random), respondents enter the sample uncontrolled — mostly out of convenience. Often simply because they are in the right place at the right time.
Judgmental Sampling
Judgmental sampling is a form of convenience sampling in which respondents enter the sample at the discretion of the researcher..
Quota Sampling
The sample is drawn according to predefined control characteristics (e.g. gender, age, income), so that it reflects the structure of the population proportionally. The objects are usually selected at random — but they must fulfill the quota plan.
| Control characteristic | Population | Sample | |
|---|---|---|---|
| Share % | Share % | Count | |
| Gender — male | 48 | 48 | 480 |
| female | 52 | 52 | 520 |
| Total | 100 | 100 | 1000 |
| Age — 18–30 | 27 | 27 | 270 |
| 31–45 | 39 | 39 | 390 |
| 45–60 | 16 | 16 | 160 |
| over 60 | 18 | 18 | 180 |
| Total | 100 | 100 | 1000 |
Often used in online surveys.
Snowball Sampling also chain sampling
Sampling
Sampling Techniques
Simple and Systematic Random Sampling
- Each element is selected independently of all others. This means:
- Each element of the population has a known and equal probability of being selected.
- Every possible sample of size n has a known probability of actually being drawn.
- First, a starting element is selected at random; then every i-th element is drawn from the sampling frame.
- The interval i results from the size of the population N relative to the size of the sample n: i = N / n
Stratified Sampling
The population is first divided into non-overlapping strata. Then a (dis-)proportional share is drawn at random from each stratum. Elements within a stratum should be similar to one another.
| Stratum | A | B | C |
|---|---|---|---|
| Population size | 100 | 200 | 300 |
| Sampling fraction | ½ | ½ | ½ |
| Sample size | 50 | 100 | 150 |
| Stratum | A | B | C |
|---|---|---|---|
| Population size | 100 | 200 | 300 |
| Sampling fraction | ⅕ | ½ | ⅓ |
| Sample size | 20 | 100 | 100 |
Cluster Samplingalso called cluster sampling
The population is divided into exclusive clusters. Then entire clusters are selected at random and enter the sample in full.
Sampling
Strengths and Weaknesses of Basic Sampling Techniques
| Technique | Strengths | Weaknesses |
|---|---|---|
| Non-probability sampling techniques | ||
| Convenience Sampling | Least expensive, least time-consuming, most convenient | Prone to error, not representative; not recommended for descriptive and causal research |
| Judgmental Sampling | Low cost, convenient, not time-consuming | Subjective, results not generalizable |
| Quota Sampling | Certain characteristics of the sample can be controlled | Prone to error, no guarantee of representativeness |
| Snowball Sampling | Enables estimation of rare characteristics | Time-consuming in fieldwork |
| Probability sampling techniques | ||
| Simple Random Sampling | Easy to understand; generalizable or representative results | Sampling frame difficult to construct, expensive, lower precision; no guarantee of representativeness |
| Systematic Sampling | Can increase representativeness; easier to implement than simple random sampling | Can decrease representativeness |
| Stratified Sampling | Includes all important subgroups of the population; high precision | Relevant stratification criteria difficult to select; multiple criteria not practical; expensive |
| Cluster Sampling | Easy to implement, cost-effective | Imprecise; complicated computation and interpretation of results |
Sampling
Determining the Sample Size
The sample size does not depend on the size of the population — it is determined by the qualitative aspects of the study:
Sample Sizes Used in Marketing Research Studies
| Type of study | Minimum size | Typical size |
|---|---|---|
| Problem identification studies (e.g. market potential) | 500 | 1.000 – 2.000 |
| Problem-solving studies (e.g. pricing) | 200 | 300 – 500 |
| Product tests | 200 | 300 – 500 |
| Test market studies | 200 | 300 – 500 |
| TV/radio/print advertising (per ad) | 150 | 200 – 300 |
| Test market audits | 10 stores | 10 – 20 stores |
| Focus groups | 6 groups | 10 – 15 groups |
A survey result — and how certain is it?
name social media as their main information channel.
But how close is this value to the true value in the population?
The answer is given by the margin of error — and it determines the required sample size.
Margin of Error Approach to Determining Sample Size
Margin of error is the measure of a survey's precision.
The smaller the margin of error, the more precise the survey's estimates are.
Margin of error approach: two formulas


Margin of Error Approach: Two Formulas
Margin of Error Approach: Two Formulas
Maximum at π = 0.5
z-Values and Maximum Margin of Error
With z = 1.96 and the maximum π = 0.5 the formula simplifies to:
… the upper bound of the error — independent of the actual proportion π.
What Is the Margin of Error?
"Which sources do you prefer to get your information from?"
Calculations show approximate values for a 95% confidence level
How Large Must the Sample Be?
The higher the desired accuracy, the larger the sample must be.
Calculations show approximate values for a 95% confidence level
What If the Population Is Small?
If the sample is larger than 10% of the population, corrections are necessary.
Otherwise the formula overestimates the required size — the finite population reduces the actual sampling error.
Calculations show approximate values for a 95% confidence level
Correcting the Sample Size
Computationally you would need n = 10,000 — with only 100 elements in the population:
= 1.000.000 / 10.099 ≈ 99
You can't survey more than the entire population — the correction brings the size down to a realistic level.
Calculations show approximate values for a 95% confidence level
Correction for a small population: ± 1 %
ncorr=n · Nn + N − 1
= 1.000.000 / 10.099
Calculations show approximate values for a 95% confidence level
Correction for a small population: ± 5 %
ncorr=n · Nn + N − 1
= 40.000 / 499
Calculations show approximate values for a 95% confidence level
Correction for a small population: ± 10 %
ncorr=n · Nn + N − 1
= 10.000 / 199
Calculations show approximate values for a 95% confidence level
Confidence Interval and Confidence Level
An estimated range of values together with the probability that this range contains the unknown parameter value.
The expected proportion of intervals that contain the parameter value across many samples.
Sample of 30 people → avg. 7.5 h. Confidence interval: 7.2 – 7.8 h (margin of error ± 0.3).
95% confidence level means: if you repeated the measurement 100 times with new samples, the true average would fall within this range in 95 of the cases.
Each bar = one confidence interval. Red = misses the true value.
Confidence Interval, Margin of Error, and Sample Size
The higher the certainty (confidence probability) we need, the wider the confidence interval becomes — and the larger the margin of error.


What the Margin of Error Formula Tells Us
The only lever to lower the margin of error is a larger n — everything else is fixed.
More certainty means a larger z → the margin of error grows → you need an even larger n to push it back down.
Data Analysis: A Concise Overview of Statistical Techniques
Types of Statistical Data Analysis
Summarizes the observations from the sample and presents them clearly.
Uses summary measures, tables, graphs and charts to describe, systematize, organize and present the collected data.
Makes statements about the generalizability of observations and conclusions from random samples to the population.
Assesses relationships between variables and quantifies them: strength and significance, predictions and estimates.
Data Analysis
Data Analysis
Frequencies and Relative Frequencies
Frequency distribution indicates, for each value, how often it occurs in the data.
Relative frequency shows the proportion (or percentage) of observations for a value.
| Favorite color | Frequency | Relative Frequency |
|---|---|---|
| blue | 10 | 10/26 ≈ 0.38 |
| red | 3 | 3/26 ≈ 0.12 |
| orange | 1 | 1/26 ≈ 0.04 |
| yellow | 3 | 3/26 ≈ 0.12 |
| green | 5 | 5/26 ≈ 0.19 |
| pink | 3 | 3/26 ≈ 0.12 |
| purple | 1 | 1/26 ≈ 0.04 |
| Total | 26 | 1.00 |
Bar Graph
Bar height = frequency or relative frequency
Bars must not touch
Pie Chart
Should always show relative frequencies.
Needs labels — directly on the chart or in the legend.
Data Analysis
Tables
| Number of Children | Frequency | Relative Frequency |
|---|---|---|
| 1 | 3 | 3/26 ≈ 0.12 |
| 2 | 8 | 8/26 ≈ 0.31 |
| 3 | 10 | 10/26 ≈ 0.38 |
| 4 | 2 | 2/26 ≈ 0.08 |
| 5 | 3 | 3/26 ≈ 0.12 |
Discrete Variable is a quantitative variable that has either a finite number of values or an infinitely countable number of values (e.g. 0, 1, 2, 3, …).
Sometimes there are too many values to create a row for each value. In that case, several values are combined into groups (classes).
| Points on the Exam | Frequency | |
|---|---|---|
| Lower Class Limit → | 50–59 | 2 |
| Upper Class Limit → | 60–69 | 5 |
| 70–79 | 7 | |
| Class Width = 90 − 80 = 10 → | 80–89 | 7 |
| 90–99 | 4 |
Tables and Histograms
| Number of Children | Frequency | Relative Frequency |
|---|---|---|
| 1 | 3 | 3/26 ≈ 0.12 |
| 2 | 8 | 8/26 ≈ 0.31 |
| 3 | 10 | 10/26 ≈ 0.38 |
| 4 | 2 | 2/26 ≈ 0.08 |
| 5 | 3 | 3/26 ≈ 0.12 |
| ∅ Time in Transit | Frequency | Relative Frequency |
|---|---|---|
| 16–17.9 | 1 | 1/15 ≈ 0.07 |
| 18–19.9 | 2 | 2/15 ≈ 0.13 |
| 20–21.9 | 1 | 1/15 ≈ 0.07 |
| 22–23.9 | 6 | 6/15 ≈ 0.40 |
| 24–25.9 | 2 | 2/15 ≈ 0.13 |
| 26–27.9 | 1 | 1/15 ≈ 0.07 |
| 28–29.9 | 1 | 1/15 ≈ 0.07 |
| 30–31.9 | 1 | 1/15 ≈ 0.07 |
Histogram
A histogram graphically depicts a grouped frequency distribution.
The height of each bar corresponds to the frequency (or relative frequency) of the class.
The widths are equal and the bars touch each other — unlike in a bar graph.
Frequency Polygon
Mark the midpoint at the top of each bar of the histogram.
Connect the midpoints with straight lines.
Bring the line back down to zero at both ends — and the polygon is complete.
Cumulative Tables and Ogives
shows the sum of the frequencies up to and including the respective row.
is the graph of the cumulative relative frequency across all classes.
| ∅ Time | Relative Frequency | Cumulative Rel. Frequency |
|---|---|---|
| 16–17.9 | 1/15 ≈ 0.07 | 0.07 |
| 18–19.9 | 2/15 ≈ 0.13 | 0.07+0.13 0.20 |
| 20–21.9 | 1/15 ≈ 0.07 | 0.20+0.07 0.27 |
| 22–23.9 | 6/15 ≈ 0.40 | 0.27+0.40 0.67 |
| 24–25.9 | 2/15 ≈ 0.13 | 0.67+0.13 0.80 |
| 26–27.9 | 1/15 ≈ 0.07 | 0.80+0.07 0.87 |
| 28–29.9 | 1/15 ≈ 0.07 | 0.87+0.07 0.94 |
| 30–31.9 | 1/15 ≈ 0.07 | 0.94+0.07 1.00 |
Data Analysis
Measures of Central Tendency

Easy to compute: just sum up and divide.
Intuitive – a single number “in the middle”; pulled up by large values and down by small ones.
Can be skewed by outliers – poor for highly variable data.
The mean of 100, 200 and −300 is 0. That is confusing.
just like the balancing point
Measures of Central Tendency
for odd nfor even nHandles outliers well – often the most accurate depiction of a group.
Splits the data into two equally sized groups.
Harder to compute: data must first be sorted.
Less well known; many confuse “median” with “average”.
of a sorted list
Measures of Central Tendency
Good for exclusive choices (this one or the other; no compromises) – works with nominal data.
Shows the choice that most people wanted (the mean often leads to a choice that no one wanted).
Easy to understand.
Requires more effort: you have to count the votes.
“The winner takes all” — there is no middle ground.
among all observations of the variable





is
Measures of Central Tendency: Using Mean and Median to Identify the Distribution Shape
Measures of Dispersion

Variance

Measures of Dispersion
Variance


The mean acts like a balance point – the average deviation from the mean is always zero.
In the variance, all deviations are squared, so that negative and positive deviations do not cancel out.

Measures of Dispersion
Deviation


Standard deviation keeps the units of measurement of the original data.
square inches
inchesRelationship between the Standard Deviation and the Shape of the Normal Distribution

Standard
Deviation
Data Analysis
Cross-Tabulations
Cross-Tabulations
Cross-tabulations summarize the joint distribution of two (or more) discrete variables in a table.
They help analyze the relationship of one variable (e.g. brand loyalty) with another (e.g. gender).
Each cell represents a combination of the categories.
Typical questions a cross-tabulation answers:
- How many brand-loyal consumers are men?
- Is the usage frequency (high, medium, low) of a product related to outdoor activities (often, sometimes, rarely, never)?
- Is familiarity with a new product related to age and level of education?
- Is ownership of a product related to income (high, medium, low)?
Cross-Tabulations
| Ownership of an expensive car | Level of education | |
|---|---|---|
| College degree | No college degree | |
| yes | 32 % | 21 % |
| no | 68 % | 79 % |
| Total | 100 % | 100 % |
| Number of cases | 250 | 750 |
Cross-Tabulations
Cross-Tabulations
relationship
| Ownership of an expensive car | High income | Low income | ||
|---|---|---|---|---|
| College degree | No college degree | College degree | No college degree | |
| yes | 20 % | 20 % | 40 % | 40 % |
| no | 80 % | 80 % | 60 % | 60 % |
| Total | 100 % | 100 % | 100 % | 100 % |
| Number of cases | 100 | 700 | 150 | 50 |
Cross-Tabulations
Relationship
| Desire for international travel | Age | |
|---|---|---|
| Under 45 | 45 and over | |
| yes | 50 % | 50 % |
| no | 50 % | 50 % |
| Total | 100 % | 100 % |
| Number of cases | 500 | 500 |
| Desire for international travel | Male | Female | ||
|---|---|---|---|---|
| < 45 | ≥ 45 | < 45 | ≥ 45 | |
| yes | 60 % | 40 % | 35 % | 65 % |
| no | 40 % | 60 % | 65 % | 35 % |
| Total | 100 % | 100 % | 100 % | 100 % |
| Number of cases | 300 | 300 | 200 | 200 |
Cross-Tabulations
Change
| Frequently go to fast-food restaurants | Family size | |
|---|---|---|
| Small | Large | |
| yes | 50 % | 50 % |
| no | 50 % | 50 % |
| Total | 100 % | 100 % |
| Number of cases | 500 | 500 |
| Frequently go to fast-food restaurants | Low income | High income | ||
|---|---|---|---|---|
| Small | Large | Small | Large | |
| yes | 50 % | 50 % | 50 % | 50 % |
| no | 50 % | 50 % | 50 % | 50 % |
| Total | 100 % | 100 % | 100 % | 100 % |
| Number of cases | 250 | 250 | 250 | 250 |
Data Analysis
Data Analysis
Hypothesis Testing
Hypothesis Testing
A five-step procedure that, based on a sample and using probability theory, determines whether a hypothesis is sufficiently supported.
In other words: a method for testing whether the results of a random sample can be generalized to the population.
A five-step approach:
- Formulating a null hypothesis and its alternative hypothesis
- Setting the significance level
- Choosing the appropriate test statistic
- Formulating the decision rule
- Calculating the metrics from the sample and making the decision
"People are mistakenly confident in their knowledge and underestimate the likelihood that their beliefs will turn out to be wrong. They tend to seek out only information that confirms what they already believe."— Max Bazerman
Hypothesis Testing
| Internet use | Gender | Total | |
|---|---|---|---|
| Male | Female | ||
| rarely | 5 | 10 | 15 |
| frequently | 10 | 5 | 15 |
| Total | 15 | 15 | n = 30 |
Hypothesis Testing


Null hypothesis (H₀) is a claim of the status quo — that there is no difference or no effect.
Alternative hypothesis (H₁) claims the opposite — that there is a difference or an effect.
Hypothesis Testing
Significance (α) — probability that a true null hypothesis is rejected.
β — probability that a false null hypothesis is accepted.
| Null hypothesis (H₀) is true | Null hypothesis (H₀) is false | |
|---|---|---|
| Null hypothesis reject | Type I errorFalse positive | Correct decisionTrue positive |
| Null hypothesis do NOT reject | Correct decisionTrue negative | Type II errorFalse negative |
Hypothesis Testing
Analogy: innocence in a criminal trial.
H₀: The defendant is innocent.
Significance (α) — probability that a true null hypothesis is rejected.
β — probability that a false null hypothesis is accepted.
| Null hypothesis (H₀) is true | Null hypothesis (H₀) is false | |
|---|---|---|
| Null hypothesis reject | Type I errorFalse positive | Correct decisionTrue positive |
| Null hypothesis do NOT reject | Correct decisionTrue negative | Type II errorFalse negative |
Hypothesis Testing
Analogy: a rustling in the bushes — is that a lion?
H₀: There is no lion in the bushes.
Significance (α) — probability that a true null hypothesis is rejected.
β — probability that a false null hypothesis is accepted.
| Null hypothesis (H₀) is true | Null hypothesis (H₀) is false | |
|---|---|---|
| Null hypothesis reject | Type I errorFalse positive | Correct decisionTrue positive |
| Null hypothesis do NOT reject | Correct decisionTrue negative | Type II errorFalse negative |
Hypothesis Testing
Significance (α) — the probability that a true null hypothesis is rejected.
β — the probability that a false null hypothesis is accepted.
Hypothesis Testing
Our example is about the distribution of non-metric variables (rare/frequent internet use; men/women) in one sample.
| Sample | Applied to | Scale level | Test statistics / Comments |
|---|---|---|---|
| One sample | Distributions | Non-metric | Kolmogorov-Smirnov and χ² test for goodness of fit; runs test for randomness; binomial test for dichotomous variables |
| Means | Metric | t-test (variance unknown); z-test (variance known) | |
| Proportions | Metric | z-test | |
| Two independent samples | Distributions | Non-metric | Kolmogorov-Smirnov test for agreement of distributions between two samples |
| Means | Metric | Two-sample t-test; F-test for equality of variances | |
| Proportions | Metric, Non-metric | z-test; χ² test | |
| Ranks / Medians | Non-metric | Mann-Whitney U-test (more sensitive than the median test) | |
| Paired samples | Means | Metric | Paired-difference t-test |
| Proportions | Non-metric | McNemar test for binary variables; χ² test | |
| Ranks / Medians | Non-metric | Wilcoxon signed-rank test (more sensitive than the sign test) |
Hypothesis Testing
One sample · distribution · non-metric → the χ² test for goodness of fit.
| Sample | Applied to | Scale level | Test statistics / Comments |
|---|---|---|---|
| One sample | Distributions | Non-metric | Kolmogorov-Smirnov and χ² test for goodness of fit; runs test for randomness; binomial test for dichotomous variables |
| Means | Metric | t-test (variance unknown); z-test (variance known) | |
| Proportions | Metric | z-test | |
| Two independent samples | Distributions | Non-metric | Kolmogorov-Smirnov test for agreement of distributions between two samples |
| Means | Metric | Two-sample t-test; F-test for equality of variances | |
| Proportions | Metric, Non-metric | z-test; χ² test | |
| Ranks / Medians | Non-metric | Mann-Whitney U-test (more sensitive than the median test) | |
| Paired samples | Means | Metric | Paired-difference t-test |
| Proportions | Non-metric | McNemar test for binary variables; χ² test | |
| Ranks / Medians | Non-metric | Wilcoxon signed-rank test (more sensitive than the sign test) |
Hypothesis Testing
The χ²test statistic (chi-square) tests the statistical significance of the relationship observed in a crosstab.
χ² tests the equality of frequency distributions — which frequencies are we comparing?
- fe — frequencies we would expect in the cells if there were no relationship.
- fo — the frequencies actually observed.
Hypothesis Testing
Hypothesis Testing

χ² should always be computed using only absolute frequencies. If the data are given in percentages (relative frequencies), they must first be converted into absolute frequencies.
Hypothesis Testing
TScal — the observed (calculated) value of the test statistic.
TScr — the critical value of the test statistic for the chosen significance level.
Hypothesis Testing
| df | 0.99 | 0.975 | 0.95 | 0.90 | 0.10 | 0.05 | 0.025 | 0.01 |
|---|---|---|---|---|---|---|---|---|
| 1 | — | 0.001 | 0.004 | 0.016 | 2.706 | 3.841 | 5.024 | 6.635 |
| 2 | 0.020 | 0.051 | 0.103 | 0.211 | 4.605 | 5.991 | 7.378 | 9.210 |
| 3 | 0.115 | 0.216 | 0.352 | 0.584 | 6.251 | 7.815 | 9.348 | 11.345 |
| 4 | 0.297 | 0.484 | 0.711 | 1.064 | 7.779 | 9.488 | 11.143 | 13.277 |
| 5 | 0.554 | 0.831 | 1.145 | 1.610 | 9.236 | 11.071 | 12.833 | 15.086 |
| 6 | 0.872 | 1.237 | 1.635 | 2.204 | 10.645 | 12.592 | 14.449 | 16.812 |
| 7 | 1.239 | 1.690 | 2.167 | 2.833 | 12.017 | 14.067 | 16.013 | 18.475 |
| 8 | 1.646 | 2.180 | 2.733 | 3.490 | 13.362 | 15.507 | 17.535 | 20.090 |
| 9 | 2.088 | 2.700 | 3.325 | 4.168 | 14.684 | 16.919 | 19.023 | 21.666 |
| 10 | 2.558 | 3.247 | 3.940 | 4.865 | 15.987 | 18.307 | 20.483 | 23.209 |
Hypothesis Testing
- H₀ — that there is no relationship — cannot be rejected.
- The relationship is 0.05 statistically not significant.
- The results observed in the sample cannot be generalized to the population.
Hypothesis Testing
| Internet use | Gender | Total | |
|---|---|---|---|
| Male | Female | ||
| rarely | 5 | 10 | 15 |
| frequently | 10 | 5 | 15 |
| Total | 15 | 15 | n = 30 |
If the sample was carefully selected and drawn, we can claim with 95% confidence that there is no such relationship.
Otherwise — we don't know.
Data Analysis
Testing the Strength of a Relationship
χ² tests only the significance of a relationship and says nothing about its strength.
Simple proof: doubling all the values in the cross-tabulation doubles χ² — but the relationship itself stays the same.
Phi Coefficient

The higher φ, the stronger the relationship between the variables.
Values > 0.30 are considered substantial.
- φ is not standardized and has an upper limit of 1 only for 2×2 tables; it depends on the table dimensions.
- φ values from different studies cannot be compared with one another.
The relationship is not particularly strong.
Contingency Coefficient

The higher C, the stronger the relationship between the variables.
Values > 0.30 are considered substantial.
Although C values have an upper limit of 1, they cannot actually reach this limit.
- C is not standardized and depends on the table dimensions.
- C values from different studies cannot be compared with one another.
The relationship is not particularly strong.
Cramer's V

The higher V, the stronger the relationship between the variables.
Values > 0.30 are considered substantial.
V values have an upper limit of 1, but they too can actually reach it only for 2×2 tables.
- V is not standardized and depends on the table dimensions.
- V values from different studies cannot be compared with one another.
The relationship is not particularly strong.
Lambda Coefficient

Indicates the extent to which knowing the value of one variable helps in predicting the other variable.
Is standardized between 0 and 1 (1 — error-free prediction, 0 — no improvement in prediction).
λ values from different studies can be compared with one another.
Knowing the gender increases prediction accuracy by a factor of 0.333, i.e. 33.3% improvement.
Lambda Coefficient
Sum of the maximum frequencies of all columns
maximum total value of a row
| Internet Use | Gender | Total (row) | |
|---|---|---|---|
| Male | Female | ||
| r = 1rarely | 5 | 10 | 15 |
| r = 2often | 10 | 5 | 15 |
| Total (column) | 15c = 1 | 15c = 2 | n = 30 |
Knowing the gender increases prediction accuracy by a factor of 0.333, i.e. 33.3% improvement.
Data Analysis
Types of Relationships between Two Variables
Unless the data come from a controlled experiment, we can only claim the existence of a relationship between the variables — but not the causal direction of that relationship.
Linear Correlation
Two variables correlate positively when higher values of one variable correspond to higher values of the other variable.
Two variables correlate negatively when higher values of one variable correspond to lower values of the other variable.
Positive Correlation
Negative Correlation
Linear Correlation Coefficient
- Values of the linear correlation coefficient always lie between −1 and 1.
- When r = +1 there is a perfect positive linear relationship between the variables.
- When r = −1 there is a perfect negative linear relationship between the variables.
- The closer r is to +1 or −1, the stronger the respective relationship.
- If r is close to 0, there is little evidence of a linear relationship — but this does not mean there is no relationship at all, just no linear.
Linear Correlation Coefficient
The (Pearson) linear correlation coefficient measures the strength of the linear relationship between two variables.

Linear Correlation Coefficient
| r-value | Interpretation |
|---|---|
| 0 to 0.3 | Very weak |
| 0.3 to 0.5 | Weak |
| 0.5 to 0.7 | Moderate |
| 0.7 to 0.9 | High |
| 0.9 to 1 | Very high |

| x | y | (xᵢ−x̄) | (yᵢ−ȳ) | (xᵢ−x̄)(yᵢ−ȳ) | (xᵢ−x̄)² | (yᵢ−ȳ)² |
|---|---|---|---|---|---|---|
| 86 | 98 | 12.5 | 13.5 | 168.75 | 156.25 | 182.25 |
| 62 | 70 | −11.5 | −14.5 | 166.75 | 132.25 | 210.25 |
| 52 | 56 | −21.5 | −28.5 | 612.75 | 462.25 | 812.25 |
| 90 | 110 | 16.5 | 25.5 | 420.75 | 272.25 | 650.25 |
| 66 | 76 | −7.5 | −8.5 | 63.75 | 56.25 | 72.25 |
| 80 | 96 | 6.5 | 11.5 | 74.75 | 42.25 | 132.25 |
| 78 | 86 | 4.5 | 1.5 | 6.75 | 20.25 | 2.25 |
| 74 | 84 | 0.5 | −0.5 | −0.25 | 0.25 | 0.25 |
| Mean | 73.5 | 84.5 | ||||
| Sum | 1514 | 1142 | 2062 |

Regression Analysis
- Can advertising spend explain changes in sales?
- Can market share be attributed to the size of the sales department?
- Is consumers' perception of quality influenced by their perception of price?
Regression Analysis
The regression analysis is a powerful and flexible tool for analyzing associative relationships between a metric dependent variable and one or more independent variables.
- determine the existence of the relationship,
- quantify the strength of the relationship,
- derive a mathematical model (formula) of the relationship,
- predict values of the dependent variable,
- account for the influence of other independent variables.
Regression Analysis
How many product units will we sell if we spend €85.000 on advertising?
| Advertising spend, €1.000 | Sales, €1.000 |
|---|---|
| 40 | 377 |
| 60 | 507 |
| 70 | 555 |
| 110 | 779 |
| 150 | 869 |
| 160 | 818 |
| 190 | 862 |
| 200 | 817 |
- Advertising spend explains 83.6% of the variance in sales.
- Every additional euro invested in advertising brings €2.82 of additional sales.
- €85.000 of advertising spend results in 2.8239 · 85 + 352.07 = 592.1 (thousand €) of sales.
Advanced Techniques of Market Analysis
Advanced Techniques of Market Analysis
Conjoint Analysis
Conjoint Analysis
The conjoint analysis is a set of techniques used in market research to analyze attribute-based preferences of consumers — i.e. to determine how important different product features and feature levels are to consumers.
What should our new product look like?
- Evaluation of holistic objects
- Decomposition of preferences
- Segmentation
- Development of new products
- Pricing

Conjoint Analysis
| not at all important | extremely important | ||||
|---|---|---|---|---|---|
| Manufacturer/Brand | |||||
| Processor performance, GHz | |||||
| Memory, GB | |||||
| Display size | |||||
| Price |
Conjoint Analysis
| not at all important | extremely important | ||||
|---|---|---|---|---|---|
| Manufacturer/Brand | |||||
| Processor performance, GHz | |||||
| Memory, GB | |||||
| Display size | |||||
| Price |
Conjoint Analysis
| Question: "Please indicate how important the following PC features are for your next purchase decision?" not at all | important extremely | ||||
|---|---|---|---|---|---|
| important | |||||
| Manufacturer/Brand | |||||
| Processor performance, GHz | |||||
| Memory, GB | |||||
| Display size |
Conjoint Analysis
The relative importance of individual product attributes can be measured more accurately when respondents evaluate them as a single stimulus (CONsidered JOINTly) than when each attribute is evaluated in isolation:
- Many consumers are unable to determine the relative importance of individual product attributes.
- Individual attributes are perceived differently in isolation than in the combination that makes up a whole product.
- Constructing the preferred combination of attributes strains respondents' cognitive abilities — “all attributes are important.”
- Social desirability and a sharply defined self-image motivate respondents to give some attributes great weight even when they play no role (e.g. environmental friendliness, wealth, reduced importance of price).
- Some respondents deliberately try to manipulate the results by giving “advantageous” answers (e.g. overstating the importance of price).
Conjoint Analysis
Advanced Techniques of Market Analysis
Market Simulations
Given the following preference values — which product should we offer on the market?
| Product alternatives | |||
|---|---|---|---|
| Blue | Red | Yellow | |
| Respondent #1 | 50 | 40 | 10 |
| Respondent #2 | 0 | 65 | 75 |
| Respondent #3 | 40 | 30 | 20 |
| Mean | 30 | 45 | 35 |
Market Simulations
Given the following preference values — which product should we offer on the market?
| Product alternatives | |||
|---|---|---|---|
| Blue | Red | Yellow | |
| Respondent #1 | 50 | 40 | 10 |
| Respondent #2 | 0 | 65 | 75 |
| Respondent #3 | 40 | 30 | 20 |
| Mean | 30 | 45 | 35 |
Market Simulations
Given the following preference values — which product should we offer on the market?
| Product alternatives | Choice | |||
|---|---|---|---|---|
| Blue | Red | Yellow | ||
| Respondent #1 | 50 | 40 | 10 | Blue |
| Respondent #2 | 0 | 65 | 75 | Yellow |
| Respondent #3 | 40 | 30 | 20 | Blue |
| Mean | 30 | 45 | 35 | Red |
Market Simulations
Suppose: 80 % of customers prefer round things, 20 % prefer square ones. What kind of things should we bring to market?
The choice seems obvious — go where most customers are:
there are already 10 competitors that ALL offer round things?
10 providers compete for the same 80% — the 20% “square” are completely unoccupied.
Why Market Simulations?
Market simulations reflect reality better than purely data-driven models — and deliver decisions that actually hold up in competition.
Closer to reality
- capture the idiosyncratic preferences of segments and individuals
- account for preferences and competing offerings on the market
Niches, not just mass
No compulsion to fixate on the “fat” part of the market — unoccupied segments can also yield good profit.
“Test laboratory”
A multitude of real market opportunities can be played through risk-free — together with their possible outcomes.
Actionable for management
The results are easy to understand and can be translated directly into decisions.
What Do Market Simulations Do?
The basic process runs individually for each respondent — and is then aggregated into market shares:
What Do Market Simulations Do?
The product choice per respondent is determined by so-called choice rules — e.g.:
First-choice rule
- The product with the highest utility is chosen.
- Selection probability = 100 % for this product, 0 % for all others.
BTL model
- Selection probability depends on the relative utility share in the market.
- Even products with low preference or utility value receive a positive probability.

Logit rule
- Selection probability increases with growing contrast in product utility.
- Enables an a-priori adjustment of simulated to real market shares.

Advanced Techniques of Market Analysis
Market Segmentation
Market segmentation refers to dividing the "relevant market" into groups of consumers who are internally homogeneous and externally heterogeneous — and forms the basis for differentiated market cultivation.
From a mixed overall group, internally homogeneous, mutually distinct segments emerge.
Developing efficient product differentiation strategies — for the most optimal exploitation of the potential of individual segments.
Effective Market Segmentation
Six criteria determine the efficiency and economic viability of a segment solution:
Identifiability
Consumers can be identified on the basis of easily measurable variables.
Substantiality
Segments must be large enough to amortize investments.
Accessibility
Targeted communication or targeted use of the marketing mix is possible.
Stability
… over the period of planning, implementation and effect of segment-specific measures.
Behavioral relevance
Uniform response to segment-specific measures (e.g. price change).
Actionability
Meaningful and helpful in formulating the marketing mix.
Typology of Segmentation Bases
- User status & usage situation
- Usage frequency & intensity
- Brand loyalty & allegiance
- Psychographic characteristics
- Values
- Personality & lifestyle
- Benefit perceptions / product benefit
- Attitudes & perception
- Preferences, motives & intentions
Benefit Segmentation
The benefit that consumers expect from products is considered one of the most relevant segmentation bases of all:
… The benefits which people are seeking in consuming a given product are the basic reasons for the existence of true market segments. Thus, benefit is the most relevant segmentation base.
… Benefit is one of the most popular bases of segmentation — for the purpose of market understanding, positioning, the development of new product concepts as well as advertising and distribution strategies. All of this owing to its actionability.
Evaluation of Segmentation Bases
| Type of base | Identifi- ability |
Substan- tiality |
Accessi- bility |
Stability | Action- ability |
Behavioral relevance |
|---|---|---|---|---|---|---|
| 1 · General, observable | ++ | ++ | ++ | ++ | – | – |
| 2 · Specific, observable | ||||||
| Purchase | + | ++ | – | + | – | + |
| Usage | + | ++ | + | + | – | + |
| 3 · General, not observable | ||||||
| Personality | ± | – | ± | ± | – | – |
| Lifestyle | ± | – | ± | ± | – | – |
| Psychographic characteristics | ± | – | ± | ± | – | – |
| 4 · Specific, not observable | ||||||
| Psychographic characteristics | ± | + | – | – | ++ | ± |
| Perception | ± | + | – | – | + | – |
| Benefit or benefit perceptions | + | + | – | + | ++ | ++ |
| Intentions | + | + | – | ± | – | ++ |
Highlighted — the most promising bases in practice: In combination they best cover the segmentation criteria.
How do the segments differ?
Benefit Segmentation Paradox
When segmenting by expected benefit, one must clearly distinguish between two types of variables:
Discriminating variables
Important for separating the sample into internally homogeneous segments.
Driver variables
Important because they represent the benefit or characteristics that respondents demand most within each segment.
Advanced Techniques of Market Analysis
Positioning
Positioning aligns all marketing activities with the preference structures of potential customers — taking into account the competing products.
Objective: to design the company's offerings so that the actual characteristics as perceived by customers are brought into alignment with the desired target characteristics.
The Role of Perception
The same respondents, the same colas — only the knowledge of the brand changes. The preference flips:


Blind test
Respondents don't know which cola they are drinking"Open" test
Respondents know which cola they are drinkingPerceptual Positioning Maps
men
special
occasion
for money
filling
women
Perceptual positioning maps depict the positions of competing products, brands or companies in a "virtual" attribute space — the way consumers perceive the entire product category.
Axes
The latent product attributes that differentiate best.
Vectors
Indicate the direction and strength of perceived product attributes.
Distances
Between two alternatives correspond to the degree of their perceived (dis)similarity.
Objective of Positioning
The objective of positioning is to occupy such a position in consumers' perception that is:
- as close to the ideal point as possible, and
- as far from the competition as possible.
Each axis is a latent attribute; each product occupies a perceived position.
Perceptual map of armchair designs
Perceptual Positioning Maps
men
special
occasion
for money
filling
women
Perceptual Positioning Maps
Competitive intensity
The closer the brands are, the more similar they are in consumers' perception — the stronger ("more direct") the competition.
men
special
occasion
for money
filling
women
Perceptual Positioning Maps
Attribute vectors
Consumers' perception of brands: The farther a brand is from the origin along an attribute vector, the more strongly this attribute is expressed in it.
men
special
occasion
for money
filling
women
Perceptual Positioning Maps
Relationships between attributes
The smaller the angle between attribute vectors, the higher their pairwise correlation.
men
special
occasion
for money
filling
women
Perceptual Positioning Maps
The length of the attribute vector indicates its degree of differentiation
The longer the vector, the more strongly this attribute differentiates between the beers.
men
special
occasion
for money
filling
women
Perceptual Positioning Maps
Axes differentiate most strongly
Axes are "virtual" attribute vectors that differentiate most strongly. Their label is usually derived from the neighboring vectors.
men
special
occasion
for money
filling
women
Perceptual Positioning Maps
Perceptual Positioning Maps and Segmentation
Straightforward assessment of segment potential and attractiveness; straightforward formulation of positioning strategies, statements and advertising campaigns.
Possible themes for advertising campaigns:
men
special
occasion
for money
filling
women
A
C
B
Example: Perception Map of Pain Relievers
Test yourself:
Example: Perception Map of Pain Relievers
Test yourself:
On Importance of Perception in Positioning
A curious but typical case
40 foods: subjective perception vs. objective attributes.
On Importance of Perception in Positioning
A curious but typical case
40 foods: subjective perception vs. objective attributes.
The perception or assessment of many food attributes often has nothing or very little to do with the actual content of these attributes.
Reporting Results
Managers should easily understand the report, trust the results and know which actions they should take.
The final stage of the research process
The report is the most frequently underestimated step of the research process.
Even a methodologically perfect study loses its value if the results are communicated poorly.
Present results so that they feed directly into the decision — not just summarize, but interpret.
Why the report and presentation matter

After the project ends, often only the written report remains — it serves as a historical record of the entire study.

The management decision rests on the report. A weak finish devalues all the preceding research.

Many managers experience only the report and the talk — and judge by them the quality of the entire project.

The decision for future research or the same service provider hinges on the perceived benefit.
From data result to follow-up
The results should serve directly as input into the decision-making; where useful, conclusions are drawn and actionable recommendations given.
Discuss the key findings, conclusions and recommendations before writing with the decision-makers — this ensures fit, acceptance and delivery dates.
The typical structure — eleven elements
- Cover letter
- Title page
- Table of contents
- Executive Summary
- Problem definition
- Approach & research design
- Data analysis
- Results
- Conclusions & recommendations
- Limitations & caveats
What comes before the report itself
Delivers the report, sums up the project experience (without results) and points to necessary follow-up steps.
Title in a manager's tone rather than "research-speak"; details on researcher and client, date.
Main and sub-headings with page numbers; then lists of tables, figures, appendices.
Briefly describes problem, approach and design and devotes one section to the key results, conclusions and recommendations. Written last — only once the entire report is finished.
The substantive core
Background, conversations with decision-makers and experts, then a clear management and research question.
Theoretical foundations, models, hypotheses; methods presented graphically and non-technically, details in the appendix.
Justify the analysis plan and techniques, explain them in simple terms with examples.
The longest part. Structured by form of analysis, data collection method or objectives — tied to the information needs.
Don't just summarize: interpret and — where possible — derive actionable recommendations.
Name limitations in a balanced way, without undermining confidence; appendix with authorization, questionnaire, sample.
Principles of good writing
Write for the decision-makers; avoid jargon, put technical terms in the appendix.
Logical structure, headings, short clear sentences; have outsiders proofread it.
Careful design; vary typography — but only as far as it supports understanding.
Present design, results and conclusions accurately, do not tailor them to the expectations of management.
Reinforce key information with tables/figures — and bring figures to life with quotes.
Leave out everything unnecessary — but never at the expense of completeness.
"The readers of your reports are busy people — hardly anyone can balance report, coffee and dictionary all at once."Guidelines for tables
- Number & Title — a unique Arabic number, a short descriptive title, referenceable in the text.
- Arrangement of Data — order by the most important aspect: time, magnitude or alphabetically.
- Unit of Measurement — state the base clearly (column vs. row percentages, sample size).
- Reading Aids — lines, shading or white space guide the eye across the row.
- Headings, Stubs, Footnotes — column heads, left margin column, explanatory footnotes.
- Source — for secondary data, name the data source.
| EMU Impact | Total | Single | Dual | Multiple |
|---|---|---|---|---|
| Existing relationships | 46 % | 41 % | 45 % | 48 % |
| Fewer banks (Eurozone) | 33 % | 31 % | 30 % | 35 % |
| One main bank coordinates | 33 % | 43 % | 39 % | 29 % |
| Fewer banks per country | 22 % | 15 % | 17 % | 25 % |
Chart types at a glance
Geographic and positioning maps show locations, customers, competitors — the basis of geodemographics.
Simple relative frequencies. Not for time series or multiple variables.
Connects data points — ideal for trends over time; multiple series comparable.
Shows absolute/relative magnitudes and differences. Histogram = vertical frequencies.
Few data points, to represent differences between groups qualitatively.
Represent process steps or the linking of qualitative ideas.
Two rules you won't forget
Suitable for simple relative frequencies — not for time series or relationships between multiple variables.
3D distorts relative magnitudes and confuses the audience. Programs offer many 3D options — hardly any presents data clearly and without distortion.
The talk shapes the first impression
Script/outline after the report, rehearse repeatedly.
Know the background, interest and stake of the listeners.
Flip chart, projector, software — never lose sight of the message.
Vary gestures, eye contact, volume and pace; a strong close.

Research Follow-up
Explain technical parts, help with implementation, discuss follow-up projects and integrate results into the MIS/DSS.
While it's still fresh, ask critically: "Could the project have been carried out more effectively or efficiently?"
The quality of the personal interaction between manager and researcher shapes the perceived quality of the report itself. Trust influences relationship quality, engagement, retention — and ultimately how strongly the market research is actually used.
Reports across countries and languages
For management in different countries and languages, separate, reader-specific versions — comparable in content, possibly different in format.
When presenting, observe cultural norms — humor is not appropriate everywhere. Adapt recommendations to be country-specific where needed.
Language versions maintained in parallel (e.g. DE / RU / KZ) are not a translation problem but a reporting problem: field-by-field localization, comparable metrics, consistent terms.
Integrity in interpretation and reporting
of the market researchers surveyed name questions of research integrity as their most difficult ethical problem.
- Ignoring relevant data
- Compromising the research design
- Deliberately misusing statistics
- Falsifying numbers or altering results
- Reinterpreting results to favor a particular view
- Withholding information
Resist the temptation: Shaping ambiguous findings into a "coherent, well-formed story" is satisfying — but unethical. Maintain objectivity, even when nothing significant emerges.
From "Push" to "Pull"
- Printed report, "push"
- Password-protected intranet reports
- First multimedia reports on the web
- Searchable, retrievable worldwide
- Interactive dashboards, "pull"
- Real-time data & live filters
- Crosstabs and weighting on demand
- Linked reports, rules for robustness
The key takeaways
The report and presentation determine the perceived value of the entire study.
Prepare results as decision input — with actionable recommendations.
Avoid jargon, structure clearly, let text and visualization reinforce each other.
≤ 7 pie segments, be careful with 3D, label tables cleanly.
The talk lives from the presenter — not from the slide.
Stay objective, follow up with the client, evaluate your own project.
About the Author

Dr. Paul Marx was a Professor of Marketing at the University of Siegen, where he conducted research particularly on e-commerce, preference measurement, new media, big data, and recommender systems. He studied aero- and hydrodynamics as well as management at Novosibirsk State Technical University (Russia) and subsequently held senior positions in marketing. In 2000 he moved to Germany, deepened his studies in economics at the University of Hannover, and founded the online service for online surveys eQuestionnaire — today QUESTIONSTAR. Paul earned his doctorate at the Bauhaus University of Weimar and published his research in leading international journals, including the Journal of Marketing.
References & License
- Backhaus, K., Erichson, B., Plinke, W., Weiber, R. (2015): "Multivariate Analysemethoden: Eine anwendungsorientierte Einführung", Springer Gabler, 14th edition.
- Brandt, D. R. (1996): "Secure Customer Index", Maritz Research.
- Bruner, G. C. (2012): "Marketing Scales Handbook", Vol. 6.
- Burke, Inc. — burke.com.
- Chuang, S.-C., Chen, H.-C. (2008): International Journal of Design.
- Malhotra, N. K. (2020): "Marketing Research: An Applied Orientation", Prentice Hall, 6th edition.
- Marx, P. (1979–2026) — own international experience.
- Moore, W. L., Pessemier, E. A. (1993), p. 145.
- Myers, J. H. (1996): "Segmentation & Positioning for Strategic Marketing Decisions", South Western Educ. Pub.
- Noelle-Neumann, E., Petersen, T. (1998): "Alle, nicht jeder. Einführung in die Methoden der Demoskopie", Springer, p. 192.
- Reichheld, F. (2003): "The One Number You Need to Grow", Harvard Business Review.
- Reichheld, F., Markey, R. (2011): "The Ultimate Question 2.0", Harvard Business Review Press.
- Sullivan III, M. (2010): "Statistics: Informed Decisions Using Data", Pearson, 3rd edition.
- Visocky O'Grady, J. & K. (2009): "The Information Design Handbook", HOW Books.
- Course "Statistics I" of Elgin Community College.
Disclaimer: The author and affiliated persons/organizations cannot be held responsible for any violation of license terms unless caused by their active conduct. Brand names and registered trademarks are the property of their respective owners and are named for descriptive purposes only. Errors excepted.
This presentation is subject to the Creative Commons Attribution-NonCommercial-ShareAlike license, unless otherwise stated. Any use or distribution requires a reference to this presentation and the explicit mention of Dr. Paul Marx and QUESTIONSTAR.