Reliability and Validity in Research
The quality criteria that make findings trustworthy, types of reliability and validity for quantitative work, and their qualitative equivalents.
Reliability and validity are the twin standards by which research quality is judged. Reliability concerns consistency. Validity concerns accuracy, whether you are measuring what you claim to measure. A methods chapter that ignores them invites the marker to distrust every finding.
Reliability: Consistency
- Test-retest: does the measure give the same result on repeat administration?
- Internal consistency: do items intended to measure one construct correlate? Reported as Cronbach's alpha (≥0.7 is a common threshold).
- Inter-rater: do independent raters/coders agree? Reported via Cohen's kappa or percentage agreement.
Validity: Accuracy
- Content validity: do the items cover the full construct?
- Construct validity: does the measure actually capture the theoretical concept?
- Criterion validity: does it correlate with an external benchmark (concurrent or predictive)?
- Internal validity: can we attribute the effect to the cause, free of confounds? (Central to experiments.)
- External validity: do the findings generalise beyond the sample and setting?
The Tension
Tightly controlled studies maximise internal validity but can weaken external validity (artificial conditions may not generalise). Naturalistic studies do the reverse. There is no free lunch. A strong methods chapter acknowledges the trade-off it has chosen and why.
Qualitative Equivalents
Reliability and validity are positivist concepts. Qualitative research uses parallel criteria from Lincoln and Guba (1985): credibility (are findings believable?, via member checking, triangulation), transferability (rich description letting readers judge applicability), dependability (a consistent, auditable process), and confirmability (findings grounded in data, not researcher bias). Use these terms in qualitative work rather than forcing validity language onto it.
Writing It Up
Do not merely define these terms, state what you did to secure them: piloting instruments, reporting alpha, using two coders, providing thick description. Threats you cannot eliminate belong in your limitations, honestly stated.
Checklist
- Have I addressed the reliability relevant to my method?
- Have I addressed the relevant types of validity?
- For qualitative work, did I use credibility/transferability/dependability/confirmability?
- Are unavoidable threats acknowledged in limitations?
- Have I reported the actual statistics rather than just naming the concept?
- Does my language match my paradigm, validity for quantitative, trustworthiness for qualitative?
How to Demonstrate Each Type in Practice
Definitions earn no marks on their own. What earns marks is evidence that you did something. This table pairs each concept with the action that demonstrates it and the thing you report.
| Concept | What you do | What you report |
|---|---|---|
| Internal consistency | Administer a multi-item scale | Cronbach's alpha, or omega, per subscale |
| Test-retest reliability | Re-administer after an interval | Correlation between the two administrations, plus the interval |
| Inter-rater reliability | Two coders rate the same material | Cohen's kappa or percentage agreement |
| Content validity | Expert panel reviews items against the construct | Who reviewed, what changed |
| Construct validity | Factor analysis, or correlate with related measures | Factor structure, convergent and discriminant correlations |
| Criterion validity | Compare against an external benchmark | Correlation with the criterion, and whether it was concurrent or predictive |
| Internal validity | Randomise, control, blind | The controls used and the confounds that remain |
| External validity | Describe sample and setting fully | Who the findings plausibly extend to |
Reading Cronbach's Alpha Honestly
Alpha is the most reported and most misread statistic in student dissertations. Tavakol and Dennick set out the two errors that matter. Alpha is not a measure of unidimensionality, and a very high value is not automatically good news.
| Value | Usual reading | What to say |
|---|---|---|
| Below 0.60 | Questionable internal consistency | Report it and treat the subscale findings cautiously |
| 0.70 to 0.90 | Generally acceptable | Report per subscale, not for the whole instrument at once |
| Above 0.95 | Possible item redundancy | Consider whether items are simply rephrasings of each other |
Two further points examiners look for. Alpha rises as you add items, so a long scale can post a comfortable figure without measuring one coherent thing. And alpha belongs to your data, not to the instrument, so report the value you obtained rather than the one printed in the original paper.
Securing Trustworthiness in Qualitative Work
The parallel criteria are often listed and rarely operationalised. Each one corresponds to something you actually do, and naming the action is what makes the claim credible.
Credibility is built by spending enough time with the data to know it, by triangulating across sources or methods, and by testing your reading against someone else's. Member checking, where you take your interpretation back to participants, is powerful but not always appropriate, since participants may not recognise an analytic reading of their own words. If you use it, say what you did with disagreement.
Transferability is not generalisation. You are not claiming your findings hold elsewhere. You are giving the reader enough detail about participants, setting and context to decide for themselves whether the findings speak to their situation. Thin description forecloses that judgement and weakens the study.
Dependability means someone could follow your decision trail. Keep a record of how coding developed, what you merged, what you discarded and why. An audit trail in an appendix is concrete evidence, and it is far more convincing than a sentence asserting that the process was rigorous.
Confirmability asks whether the findings come from the data rather than your assumptions. A reflexive statement is the usual vehicle. Say who you are in relation to the topic, what you expected to find, and how that might have shaped what you noticed. This is not a confession, it is methodological transparency, and examiners read its absence as a gap.
Common Threats and What Blunts Them
| Threat | What goes wrong | Mitigation |
|---|---|---|
| Selection bias | Groups differ before the intervention | Randomise, or measure and adjust for baseline differences |
| Maturation | Participants change over time regardless | Include a control group |
| Testing effects | The first test changes later performance | Use parallel forms, or lengthen the interval |
| Attrition | Dropouts differ systematically from completers | Report the rate and compare leavers with stayers |
| Social desirability | Participants answer as they think they should | Anonymise responses and say so in the instructions |
| Researcher bias | Expectation shapes coding or interpretation | Blind coding, second coder, reflexive statement |
Common Mistakes and Fixes
| Mistake | Fix |
|---|---|
| Defining the terms without applying them | State what you did, per concept, in the methods chapter |
| Using validity language in a qualitative study | Switch to credibility, transferability, dependability, confirmability |
| One alpha for a multi-subscale instrument | Report alpha separately for each subscale |
| Quoting the original paper's reliability figure | Report the figure from your own sample |
| Claiming generalisability from a convenience sample | Limit the claim to the population you actually sampled |
| Hiding a weak result in the appendix | Report it in the body and discuss what it means |
Discipline Notes
- Psychology: the fullest treatment, with construct validity and factor structure often expected in detail.
- Health and nursing: reliability of instruments and inter-rater agreement carry particular weight.
- Education: mixed designs are common, so you may need both vocabularies in one chapter, clearly separated.
- Business and management: survey construct validity matters, and reviewers often ask about common method bias.
- Sociology and anthropology: trustworthiness criteria usually replace validity language entirely.
Where Ethical Support Fits
Statistical advice is a normal part of research training, and most universities employ someone whose job is exactly this. Asking how to interpret an alpha value, or whether a second coder is expected in your field, is legitimate. So is having a methods chapter proofread for clarity.
What must stay yours is the data, the analysis you ran, and the honest reporting of what it showed. A weak reliability figure is a finding to discuss, not a number to adjust.
Frequently Asked Questions
Can something be reliable but not valid?
Yes, and it is the standard illustration. A scale that reads three kilograms heavy gives the same wrong answer every time. It is perfectly reliable and completely invalid. The reverse is not possible, because a measure that varies randomly cannot be measuring the construct consistently.
What alpha do I need?
There is no universal threshold, though 0.70 is the figure most often cited for research purposes. Report what you obtained and discuss it rather than treating any number as a pass mark.
Do I need reliability statistics for a qualitative study?
Usually not. Reflexive thematic analysis in particular treats coding as interpretation rather than measurement, so inter-rater reliability can be conceptually inappropriate. Use trustworthiness criteria instead, and say why.
My instrument is validated already. Do I still report anything?
Yes. Cite the validation evidence, then report internal consistency in your own sample. A validated instrument can behave differently in a new population or a new language.
How do I improve external validity in a student project?
Often you cannot, given time and access. What you can do is describe your sample and setting in enough detail that a reader can judge where the findings might transfer, then state the boundary explicitly in your limitations.
Where do these go in the dissertation?
Methods, mostly, as part of justifying your instruments and procedures. Residual threats belong in limitations, and any that shape how you read your results belong in the discussion too.
Your Next Step Today
Open your methods chapter and find every sentence that names reliability or validity. If a sentence defines the term without saying what you did about it, rewrite it as an action plus a reported result. That single pass is usually what moves the section from descriptive to assessed.
Trusted Sources
- Tavakol M and Dennick R, Making sense of Cronbach's alpha, International Journal of Medical Education 2011. Open access. Accessed 11 August 2026.
- APA Dictionary of Psychology, construct validity. Accessed 11 August 2026.
- APA Dictionary of Psychology, test-retest reliability. Accessed 11 August 2026.
- University of Southern California, The Methodology, in Organizing Your Social Sciences Research Paper. Accessed 11 August 2026.
- University of Southern California, Limitations of the Study. Accessed 11 August 2026.
- Lincoln YS and Guba EG, Naturalistic Inquiry (Sage 1985), the source of the trustworthiness criteria.
Terminology and expectations vary by discipline. Where this guide and your department's handbook differ, follow your department.
Related services
Related guides
- Writing the Methodology Chapter: Design, Sampling and Justification
- Sampling Methods in Research Explained
- Choosing the Right Statistical Test
- Mixed-Methods Research Design Explained
All academic writing guides or browse our academic writing services.