Reliability and Validity in Research

The quality criteria that make findings trustworthy, types of reliability and validity for quantitative work, and their qualitative equivalents.

Reliability and validity are the twin standards by which research quality is judged. Reliability concerns consistency. Validity concerns accuracy, whether you are measuring what you claim to measure. A methods chapter that ignores them invites the marker to distrust every finding.

Reliability: Consistency

Validity: Accuracy

The Tension

Tightly controlled studies maximise internal validity but can weaken external validity (artificial conditions may not generalise). Naturalistic studies do the reverse. There is no free lunch. A strong methods chapter acknowledges the trade-off it has chosen and why.

Qualitative Equivalents

Reliability and validity are positivist concepts. Qualitative research uses parallel criteria from Lincoln and Guba (1985): credibility (are findings believable?, via member checking, triangulation), transferability (rich description letting readers judge applicability), dependability (a consistent, auditable process), and confirmability (findings grounded in data, not researcher bias). Use these terms in qualitative work rather than forcing validity language onto it.

Writing It Up

Do not merely define these terms, state what you did to secure them: piloting instruments, reporting alpha, using two coders, providing thick description. Threats you cannot eliminate belong in your limitations, honestly stated.

Checklist

How to Demonstrate Each Type in Practice

Definitions earn no marks on their own. What earns marks is evidence that you did something. This table pairs each concept with the action that demonstrates it and the thing you report.

ConceptWhat you doWhat you report
Internal consistencyAdminister a multi-item scaleCronbach's alpha, or omega, per subscale
Test-retest reliabilityRe-administer after an intervalCorrelation between the two administrations, plus the interval
Inter-rater reliabilityTwo coders rate the same materialCohen's kappa or percentage agreement
Content validityExpert panel reviews items against the constructWho reviewed, what changed
Construct validityFactor analysis, or correlate with related measuresFactor structure, convergent and discriminant correlations
Criterion validityCompare against an external benchmarkCorrelation with the criterion, and whether it was concurrent or predictive
Internal validityRandomise, control, blindThe controls used and the confounds that remain
External validityDescribe sample and setting fullyWho the findings plausibly extend to

Reading Cronbach's Alpha Honestly

Alpha is the most reported and most misread statistic in student dissertations. Tavakol and Dennick set out the two errors that matter. Alpha is not a measure of unidimensionality, and a very high value is not automatically good news.

ValueUsual readingWhat to say
Below 0.60Questionable internal consistencyReport it and treat the subscale findings cautiously
0.70 to 0.90Generally acceptableReport per subscale, not for the whole instrument at once
Above 0.95Possible item redundancyConsider whether items are simply rephrasings of each other

Two further points examiners look for. Alpha rises as you add items, so a long scale can post a comfortable figure without measuring one coherent thing. And alpha belongs to your data, not to the instrument, so report the value you obtained rather than the one printed in the original paper.

Securing Trustworthiness in Qualitative Work

The parallel criteria are often listed and rarely operationalised. Each one corresponds to something you actually do, and naming the action is what makes the claim credible.

Credibility is built by spending enough time with the data to know it, by triangulating across sources or methods, and by testing your reading against someone else's. Member checking, where you take your interpretation back to participants, is powerful but not always appropriate, since participants may not recognise an analytic reading of their own words. If you use it, say what you did with disagreement.

Transferability is not generalisation. You are not claiming your findings hold elsewhere. You are giving the reader enough detail about participants, setting and context to decide for themselves whether the findings speak to their situation. Thin description forecloses that judgement and weakens the study.

Dependability means someone could follow your decision trail. Keep a record of how coding developed, what you merged, what you discarded and why. An audit trail in an appendix is concrete evidence, and it is far more convincing than a sentence asserting that the process was rigorous.

Confirmability asks whether the findings come from the data rather than your assumptions. A reflexive statement is the usual vehicle. Say who you are in relation to the topic, what you expected to find, and how that might have shaped what you noticed. This is not a confession, it is methodological transparency, and examiners read its absence as a gap.

Common Threats and What Blunts Them

ThreatWhat goes wrongMitigation
Selection biasGroups differ before the interventionRandomise, or measure and adjust for baseline differences
MaturationParticipants change over time regardlessInclude a control group
Testing effectsThe first test changes later performanceUse parallel forms, or lengthen the interval
AttritionDropouts differ systematically from completersReport the rate and compare leavers with stayers
Social desirabilityParticipants answer as they think they shouldAnonymise responses and say so in the instructions
Researcher biasExpectation shapes coding or interpretationBlind coding, second coder, reflexive statement

Common Mistakes and Fixes

MistakeFix
Defining the terms without applying themState what you did, per concept, in the methods chapter
Using validity language in a qualitative studySwitch to credibility, transferability, dependability, confirmability
One alpha for a multi-subscale instrumentReport alpha separately for each subscale
Quoting the original paper's reliability figureReport the figure from your own sample
Claiming generalisability from a convenience sampleLimit the claim to the population you actually sampled
Hiding a weak result in the appendixReport it in the body and discuss what it means

Discipline Notes

Where Ethical Support Fits

Statistical advice is a normal part of research training, and most universities employ someone whose job is exactly this. Asking how to interpret an alpha value, or whether a second coder is expected in your field, is legitimate. So is having a methods chapter proofread for clarity.

What must stay yours is the data, the analysis you ran, and the honest reporting of what it showed. A weak reliability figure is a finding to discuss, not a number to adjust.

Frequently Asked Questions

Can something be reliable but not valid?

Yes, and it is the standard illustration. A scale that reads three kilograms heavy gives the same wrong answer every time. It is perfectly reliable and completely invalid. The reverse is not possible, because a measure that varies randomly cannot be measuring the construct consistently.

What alpha do I need?

There is no universal threshold, though 0.70 is the figure most often cited for research purposes. Report what you obtained and discuss it rather than treating any number as a pass mark.

Do I need reliability statistics for a qualitative study?

Usually not. Reflexive thematic analysis in particular treats coding as interpretation rather than measurement, so inter-rater reliability can be conceptually inappropriate. Use trustworthiness criteria instead, and say why.

My instrument is validated already. Do I still report anything?

Yes. Cite the validation evidence, then report internal consistency in your own sample. A validated instrument can behave differently in a new population or a new language.

How do I improve external validity in a student project?

Often you cannot, given time and access. What you can do is describe your sample and setting in enough detail that a reader can judge where the findings might transfer, then state the boundary explicitly in your limitations.

Where do these go in the dissertation?

Methods, mostly, as part of justifying your instruments and procedures. Residual threats belong in limitations, and any that shape how you read your results belong in the discussion too.

Your Next Step Today

Open your methods chapter and find every sentence that names reliability or validity. If a sentence defines the term without saying what you did about it, rewrite it as an action plus a reported result. That single pass is usually what moves the section from descriptive to assessed.

Trusted Sources

Terminology and expectations vary by discipline. Where this guide and your department's handbook differ, follow your department.

Related services

Related guides

All academic writing guides or browse our academic writing services.