Saif Ali AlghamdiTransformation & Growth Advisor
تواصل
Foundation of Business Researchأسس بحوث الأعمالResearch methodologyمنهجية البحث
RESEARCH METHODOLOGY · PhDمنهجية البحث · دكتوراه

Validity and Reliabilityالصدق والثبات

SectionالقسمResearch methodologyمنهجية البحث
Reading timeزمن القراءة11 min١١ دقيقة
ByإعدادSaif Alghamdiسيف الغامدي
One

Overview

Topic: Validity, reliability, and their qualitative equivalents
Covers: The five validities, reliability tests, threats, and the four qualitative criteria
Use: Demonstrating that your findings mean what you say they mean
Level: Doctoral business research
By: Saif Alghamdi

Validity asks whether you are measuring the thing you claim to measure. Reliability asks whether the measurement would repeat. They are the two properties on which every quantitative finding depends, and the relationship between them is asymmetric: a measure can be reliable and completely invalid, but it cannot be valid without being reliable.

A bathroom scale reading three kilograms heavy every time is perfectly reliable and systematically wrong. A scale giving a different answer each time cannot be accurate even by chance in any useful sense. That asymmetry explains why reliability is reported first and why reporting it is not enough: an alpha coefficient tells a reader your items hang together, not that they hang together around the right construct.

Qualitative research faces the same underlying concern in different terms, because the question of whether an account can be trusted does not disappear when the data are words. The conventional criteria are credibility, transferability, dependability, and confirmability, and they are not softer versions of the quantitative terms but a distinct set adapted to a different logic of inquiry.

This framework covers the five kinds of validity and what evidence each requires, reliability and how it is tested, the specific threats that recur in business research, the four qualitative criteria and the procedures that support them, and how to report all of this without either inflating or apologising. It is my own synthesis, written in my own words and grounded in recognized scholarship.

Note: Validity is a property of an inference, not of an instrument. A scale is not valid in the abstract; it is valid for a particular use, with a particular population, for a particular claim. Saying a validated instrument was used answers nothing on its own.
الأول

نظرة عامة

الموضوع: الصدق والثبات ومكافئاتهما النوعية
يغطّي: أنواع الصدق الخمسة، واختبارات الثبات، والتهديدات، والمعايير النوعية الأربعة
الاستخدام: إظهار أن نتائجك تعني ما تقول إنها تعنيه
المستوى: بحوث الأعمال لمرحلة الدكتوراه
إعداد: سيف الغامدي

الصدق يسأل هل تقيس الشيء الذي تدّعي قياسه. والثبات يسأل هل يتكرر القياس. وهما الخاصيتان اللتان تعتمد عليهما كل نتيجة كمية، والعلاقة بينهما لاتماثلية: فالمقياس قد يكون ثابتًا وغير صادق بالمرة، لكنه لا يمكن أن يكون صادقًا دون أن يكون ثابتًا.

فميزانٌ يقرأ ثلاثة كيلوغرامات زيادةً في كل مرة ثابتٌ تمامًا وخاطئٌ منهجيًا. وميزانٌ يعطي إجابةً مختلفة كل مرة لا يمكن أن يكون دقيقًا ولو مصادفةً بأي معنى نافع. وذلك اللاتماثل يفسّر لماذا يُبلَّغ عن الثبات أولًا ولماذا لا يكفي الإبلاغ عنه: فمعامل ألفا يخبر القارئ أن بنودك تتماسك، لا أنها تتماسك حول البناء الصحيح.

والبحث النوعي يواجه الهاجس الكامن نفسه بعبارات مختلفة، لأن سؤال هل يمكن الوثوق بعرضٍ لا يختفي حين تكون البيانات كلمات. والمعايير المتعارَفة المصداقية وقابلية النقل والاعتمادية وإمكانية التأكيد، وهي ليست نسخًا أرخى من المصطلحات الكمية بل مجموعة متمايزة مكيَّفة على منطق تقصٍّ مختلف.

ويغطي هذا الإطار أنواع الصدق الخمسة وأي دليل يتطلبه كلٌّ، والثبات وكيف يُختبَر، والتهديدات المحددة المتكررة في بحوث الأعمال، والمعايير النوعية الأربعة والإجراءات التي تسندها، وكيف يُبلَّغ عن هذا كله بلا تضخيم ولا اعتذار. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.

ملاحظة: الصدق خاصيةٌ لاستدلال لا لأداة. فالمقياس ليس صادقًا في المطلق؛ بل صادقٌ لاستخدام بعينه، مع مجتمع بعينه، لادّعاء بعينه. وقولُ «استُخدمت أداة مصدَّقة» لا يجيب شيئًا وحده.
Two

The Five Kinds of Validity

Three concern whether the instrument measures the construct, and two concern whether the study's inferences hold. All five should be addressed, though not all at equal length.

1Contentdo the items cover the whole construct and nothing else?2Constructdoes the measure behave as the theory says it should?3Criteriondoes it predict or agree with an external standard?4Internalis the causal claim protected from alternative explanations?5Externaldoes the finding hold beyond this sample and setting?five validities; the first three concern measurement, the last two concern design

Content validity asks whether the items cover the whole construct and nothing outside it. It is established by judgement rather than by statistics: the construct definition is compared with the item pool, and subject experts review whether anything essential is missing or anything irrelevant is included. Report it by saying who reviewed the instrument, what they were asked, and what changed. A measure of managerial support consisting only of items about resources has a content validity problem no reliability coefficient will reveal.

Construct validity asks whether the measure behaves as the theory predicts. It has two halves that are tested together. Convergent validity means the measure correlates with other measures of the same or closely related constructs, and is commonly evidenced by the average variance extracted exceeding a conventional threshold. Discriminant validity means it does not correlate too highly with measures of different constructs, which is what distinguishes two variables from one variable measured twice. In business research this is the most frequently neglected evidence and the most frequently requested by reviewers.

Criterion validity asks whether the measure agrees with or predicts an external standard: does an engagement measure predict actual turnover, does a credit score predict default. Where a criterion exists this is the strongest measurement evidence available, and where none exists, which is common, say so rather than omitting the heading.

Internal validity concerns the design rather than the instrument, and asks whether the causal inference is protected from alternative explanations. It is the property experiments are built to secure and the one cross-sectional surveys cannot have. Reverse causation, omitted variables, and selection into groups are the standard threats, and a study using a design that cannot rule them out should say which ones remain open.

External validity asks how far the finding extends beyond the sample, setting, and period studied. It is bounded by the sampling and by the context, and the honest treatment is to describe the conditions under which the finding was produced in enough detail that a reader can judge transfer themselves, rather than asserting generalisability.

Note: Address all five under labelled headings, even where the answer is short. A heading reading criterion validity followed by one sentence explaining that no external criterion was available is far better than the absence of the heading, which reads as an omission rather than a decision.
الثاني

أنواع الصدق الخمسة

ثلاثة تخص هل تقيس الأداة البناء، واثنان يخصان هل تصمد استدلالات الدراسة. وينبغي معالجة الخمسة كلها، وإن لم يكن بطول متساوٍ.

١المحتوىهل تغطي البنود البناء كله ولا شيء سواه؟٢البناءهل يسلك المقياس كما تقول النظرية إنه ينبغي؟٣المحكهل يتنبأ بمعيار خارجي أو يتفق معه؟٤الداخليهل الادّعاء السببي محميّ من التفسيرات البديلة؟٥الخارجيهل تصح النتيجة خارج هذه العيّنة وهذا الموقع؟خمسة أنواع صدق؛ الثلاثة الأولى عن القياس والأخيران عن التصميم

وصدق المحتوى يسأل هل تغطي البنود البناء كله ولا شيء خارجه. ويُثبَت بالحكم لا بالإحصاء: فيُقارَن تعريف البناء ببركة البنود، ويراجع خبراء الموضوع هل نقص شيء جوهري أو أُدرج شيء لا صلة له. وأبلغ عنه بقول من راجع الأداة وما طُلب منهم وما تغيّر. فمقياسٌ لدعم المدير يتألف من بنود عن الموارد فقط فيه مشكلة صدق محتوى لن يكشفها أي معامل ثبات.

وصدق البناء يسأل هل يسلك المقياس كما تتنبأ النظرية. وله نصفان يُختبَران معًا. الصدق التقاربي يعني أن المقياس يرتبط بمقاييس أخرى للبناء نفسه أو لبناءات قريبة، ويُستدَل عليه شائعًا بتجاوز متوسط التباين المستخلَص عتبةً متعارَفة. والصدق التمييزي يعني أنه لا يرتبط بدرجة عالية جدًا بمقاييس بناءات مختلفة، وهذا ما يميّز متغيرين عن متغير واحد قيس مرتين. وفي بحوث الأعمال هذا أكثر الأدلة إهمالًا وأكثرها طلبًا من المحكّمين.

وصدق المحك يسأل هل يتفق المقياس مع معيار خارجي أو يتنبأ به: هل يتنبأ مقياس اندماج بالدوران الفعلي، وهل تتنبأ درجة ائتمان بالتعثر. وحيث يوجد محك يكون هذا أقوى دليل قياس متاح، وحيث لا يوجد، وهذا شائع، فقُل ذلك بدل إغفال العنوان.

والصدق الداخلي يخص التصميم لا الأداة، ويسأل هل الاستدلال السببي محميّ من التفسيرات البديلة. وهو الخاصية التي تُبنى التجارب لتأمينها والتي لا تستطيع المسوح المقطعية امتلاكها. والسببية العكسية والمتغيرات المُغفَلة والانتقاء إلى المجموعات تهديداتٌ معيارية، والدراسة التي تستخدم تصميمًا لا يستطيع استبعادها ينبغي أن تقول أيّها يبقى مفتوحًا.

والصدق الخارجي يسأل إلى أي مدى تمتد النتيجة خارج العيّنة والموقع والحقبة المدروسة. وتحدّه المعاينةُ والسياقُ، والمعالجة الصادقة أن تصف الشروط التي أُنتجت فيها النتيجة بتفصيل يكفي ليحكم القارئ على النقل بنفسه، لا أن تجزم بالقابلية للتعميم.

ملاحظة: عالج الخمسة كلها تحت عناوين مُسمّاة، حتى حيث تكون الإجابة قصيرة. فعنوانٌ يقرأ «صدق المحك» يتبعه جملةٌ تشرح أنه لا محك خارجي متاح خيرٌ بكثير من غياب العنوان، الذي يُقرأ إغفالًا لا قرارًا.
Three

Reliability and How It Is Tested

Reliability is consistency, and it is tested in four ways depending on what kind of consistency is in question.

Validityare you measuring the thing you say you aremeasuring?Reliabilitywould the same measurement repeat if you didit again?consistency is necessary for accuracy but does not produce itreliable and invalid is possible; valid and unreliable is not

Internal consistency asks whether the items in a scale measure the same thing. It is the reliability almost every business thesis reports, usually as a coefficient alpha, with values above 0.70 conventionally acceptable and above 0.80 good. Three cautions apply. Alpha rises with the number of items, so a long scale can look reliable while measuring several things. Alpha above about 0.95 suggests redundancy, meaning items that repeat each other. And alpha assumes the items are equally related to the construct, which is why composite reliability is increasingly reported alongside it.

Test-retest reliability asks whether the same respondents give the same answers on a second occasion. It is the appropriate test for a construct that should be stable over short periods, and it requires a second administration, typically two to four weeks apart. It is rarely done in doctoral work and is worth doing on a subsample when the construct's stability is itself in question.

Inter-rater reliability asks whether two people applying the same coding scheme reach the same result. It applies to content analysis, to observational coding, and to qualitative coding where a second coder is available. Report the statistic and the proportion of material double-coded, and describe how disagreements were resolved.

Parallel forms asks whether two versions of an instrument produce equivalent results, and matters mainly when a shortened version is used or when translation has produced two language versions that will be pooled.

Three reporting rules make the section credible. Report reliability for each scale separately, not one figure for the whole instrument, since instruments measure several constructs and a single coefficient conceals a weak one. Report it from your own data, not from the source study; the prior figure establishes that the scale can be reliable and yours establishes that it was in your population. And where a coefficient falls below threshold, report it and address it, either by identifying the item whose removal improves it, or by retaining the scale and noting the limitation, but never by omitting the number.

Note: A high alpha is not evidence of validity and is frequently mistaken for it. Five items that all ask about pay satisfaction in slightly different words will be highly reliable as a measure of job satisfaction and highly invalid as one.
الثالث

الثبات وكيف يُختبَر

الثبات اتساق، ويُختبَر بأربع طرق بحسب أي نوع من الاتساق موضع السؤال.

الصدقهل تقيس الشيء الذي تقول إنك تقيسه؟الثباتهل يتكرر القياس نفسه لو أعدته؟الاتساق لازم للدقة لكنه لا يُنتجهاالثابت غير الصادق ممكن؛ والصادق غير الثابت غير ممكن

والاتساق الداخلي يسأل هل تقيس بنود مقياسٍ الشيء نفسه. وهو الثبات الذي تُبلغ عنه كل أطروحة أعمال تقريبًا، معاملَ ألفا عادةً، والقيم فوق 0.70 مقبولة عرفًا وفوق 0.80 جيدة. وثلاثة تحذيرات تنطبق. فألفا يرتفع بعدد البنود، فقد يبدو مقياسٌ طويل ثابتًا وهو يقيس أشياء عدة. وألفا فوق نحو 0.95 يوحي بالتكرار، أي ببنود تعيد بعضها. وألفا يفترض أن البنود مرتبطة بالبناء بالقدر نفسه، ولهذا يُبلَّغ عن الثبات المركّب إلى جانبه بازدياد.

وثبات الإعادة يسأل هل يعطي المستجيبون أنفسهم الإجابات نفسها في مناسبة ثانية. وهو الاختبار الملائم لبناء ينبغي أن يكون مستقرًا عبر فترات قصيرة، ويتطلب تطبيقًا ثانيًا، بفارق أسبوعين إلى أربعة عادةً. وهو نادرٌ في عمل الدكتوراه ويستحق الإجراء على عيّنة فرعية حين يكون استقرار البناء نفسه موضع سؤال.

وثبات المرمّزين يسأل هل يصل شخصان يطبّقان مخطط الترميز نفسه إلى النتيجة نفسها. وينطبق على تحليل المحتوى وعلى الترميز الملاحظي وعلى الترميز النوعي حيث يتوفر مرمّز ثانٍ. أبلغ بالإحصاءة وبنسبة المادة المرمَّزة مزدوجًا، وصِف كيف حُلّت الاختلافات.

والصيغ المتوازية تسأل هل تُنتج نسختان من أداة نتائج مكافئة، وتهم أساسًا حين تُستخدَم نسخة مختصرة أو حين تُنتج الترجمةُ نسختين لغويتين ستُجمَعان.

وثلاث قواعد إبلاغ تجعل القسم موثوقًا. أبلغ بالثبات لكل مقياس منفصلًا، لا رقمًا واحدًا للأداة كلها، فالأدوات تقيس بناءات عدة والمعامل الواحد يخفي ضعيفًا. وأبلغ به من بياناتك أنت، لا من الدراسة المصدر؛ فالرقم السابق يثبت أن المقياس يمكن أن يكون ثابتًا ورقمُك يثبت أنه كان كذلك في مجتمعك. وحيث يقع معامل تحت العتبة، أبلغ به وعالجه، إما بتحديد البند الذي يحسّنه حذفُه، أو بالإبقاء على المقياس وتدوين الحد، لكن لا بإغفال الرقم أبدًا.

ملاحظة: ألفا العالي ليس دليل صدق ويُخلَط به كثيرًا. فخمسة بنود كلها تسأل عن الرضا عن الأجر بكلمات مختلفة قليلًا ستكون عالية الثبات مقياسًا للرضا الوظيفي وعالية اللاصدق مقياسًا له.
Four

Threats That Recur in Business Research

Six threats account for most validity problems in business studies. Each has a recognisable signature and a standard mitigation that costs little if planned in advance.

ThreatWhat it doesMitigation
Common method biasInflates correlations when all variables come from one sourceSeparate sources or times; anonymity; a marker variable
Social desirabilityRespondents report what is approved rather than what is trueAnonymity, indirect wording, behavioural rather than attitude items
Non-response biasThose who answer differ from those who do notCompare respondents with the frame; early versus late analysis
Reverse causationThe outcome may be producing the predictorTime-lagged measurement of predictor and outcome
Omitted variablesA third factor produces both observed variablesMeasure and control plausible confounds; theorise them first
Construct contaminationThe measure captures more than the constructDiscriminant validity evidence; expert review of items

Common method bias deserves the most attention because it is nearly universal in business surveys and is routinely ignored. When the same respondent reports both the predictor and the outcome at the same moment on the same instrument, part of the observed correlation comes from the shared method rather than from the relationship. Mitigations are cheap if planned: collect the outcome from a different source such as records or a supervisor, separate the two measurements in time, guarantee anonymity, and vary the response formats. Post hoc statistical tests exist and are weaker than any of these design remedies.

Social desirability is severe for anything involving ethics, discrimination, performance, or compliance. The strongest remedy is anonymity that respondents believe, which depends on how the study was introduced and by whom. Asking about behaviour rather than attitudes helps: how many times did you do X in the last month is harder to inflate than do you value X.

Reverse causation is the threat most often overlooked in the writing rather than the design. Satisfied employees perform better and better performers become satisfied are equally consistent with a cross-sectional correlation, and the reader will notice even if the author did not. Where the design cannot settle it, saying so directly is stronger than a hedge.

Address the relevant threats in a short subsection of the methodology rather than in the limitations at the end. A threat named in the methodology alongside its mitigation reads as design; the same threat appearing only in the limitations reads as something you noticed afterwards.

Note: Choose which threats matter for your design and address those, rather than listing every threat in the textbook. Four threats addressed specifically is more convincing than twelve mentioned generically.
الرابع

تهديدات متكررة في بحوث الأعمال

ستة تهديدات تفسّر معظم مشكلات الصدق في دراسات الأعمال. ولكلٍّ بصمةٌ معروفة وتخفيفٌ معياري يكلّف قليلًا إن خُطِّط له سلفًا.

التهديدما يفعلهالتخفيف
تحيّز المنهج المشتركيضخّم الارتباطات حين تأتي كل المتغيرات من مصدر واحدمصادر أو أوقات منفصلة؛ وإخفاء الهوية؛ ومتغير علامة
المرغوبية الاجتماعيةالمستجيبون يُبلغون بما يُستحسَن لا بما هو صحيحإخفاء الهوية، وصياغة غير مباشرة، وبنود سلوكية لا اتجاهية
تحيّز عدم الاستجابةمن يجيبون يختلفون عمّن لا يجيبونقارن المستجيبين بالإطار؛ وتحليل المبكر مقابل المتأخر
السببية العكسيةقد يكون المخرَج هو ما يُنتج المتنبئقياس متباعد زمنيًا للمتنبئ والمخرَج
المتغيرات المُغفَلةعاملٌ ثالث يُنتج المتغيرين الملحوظينقِس واضبط المُربِكات المعقولة؛ ونظّر لها أولًا
تلوّث البناءالمقياس يلتقط أكثر من البناءدليل صدق تمييزي؛ ومراجعة خبراء للبنود

وتحيّز المنهج المشترك يستحق أكبر انتباه لأنه شبه كوني في مسوح الأعمال ويُتجاهَل روتينيًا. فحين يُبلغ المستجيب نفسه بالمتنبئ والمخرَج في اللحظة نفسها على الأداة نفسها، يأتي جزء من الارتباط الملحوظ من المنهج المشترك لا من العلاقة. والتخفيفات رخيصة إن خُطِّط لها: اجمع المخرَج من مصدر مختلف كالسجلات أو المشرف، وافصل القياسين زمنيًا، واضمن إخفاء الهوية، ونوّع صيغ الاستجابة. والاختبارات الإحصائية اللاحقة موجودة وهي أضعف من أي من هذه العلاجات التصميمية.

والمرغوبية الاجتماعية شديدة لأي شيء يتعلق بالأخلاقيات أو التمييز أو الأداء أو الامتثال. وأقوى علاج إخفاءُ هوية يصدّقه المستجيبون، وهذا يعتمد على كيف قُدِّمت الدراسة ومن قدّمها. والسؤال عن السلوك لا عن الاتجاهات يساعد: فـ«كم مرة فعلت س في الشهر الماضي» أصعب تضخيمًا من «هل تقدّر س».

والسببية العكسية أكثر التهديدات إغفالًا في الكتابة لا في التصميم. فـ«الموظفون الراضون يؤدون أفضل» و«الأفضل أداءً يصيرون راضين» متسقان بالقدر نفسه مع ارتباط مقطعي، والقارئ سيلاحظ حتى لو لم يلاحظ المؤلف. وحيث لا يستطيع التصميم حسمه، فقولُ ذلك مباشرةً أقوى من التحويط.

وعالج التهديدات ذات الصلة في قسم فرعي قصير من المنهجية لا في الحدود في النهاية. فالتهديد المسمّى في المنهجية إلى جانب تخفيفه يُقرأ تصميمًا؛ والتهديد نفسه ظاهرًا في الحدود وحدها يُقرأ شيئًا لاحظته بعد فوات الأوان.

ملاحظة: اختر أي التهديدات تهم تصميمك وعالجها، بدل سرد كل تهديد في الكتاب المدرسي. فأربعة تهديدات معالَجة تحديدًا أقنعُ من اثني عشر مذكورًا عمومًا.
Five

The Four Qualitative Criteria

Qualitative research does not claim that a different researcher would produce identical findings, so a different vocabulary is needed. The four criteria are demanding rather than lenient, and each is supported by a specific procedure.

CredibilityTransferabilityDependabilityConfirmabilityare the findingsbelievable to those inthe setting?can a reader judgewhether it applieselsewhere?is the processtraceable andconsistent?are the findingsgrounded in data, not inthe researcher?member checking,triangulationthick description ofcontextan audit trail ofdecisionsreflexivity and rawdata extractsthe qualitative equivalents; different words for the same underlying concern

Credibility asks whether the findings are a believable account of the participants' reality. Three procedures support it. Triangulation compares accounts across sources, methods, or investigators, and its value is not agreement but the explanation of disagreement. Member checking returns the interpretation to participants and asks whether it recognises their experience, which is powerful and must be handled carefully, since participants may disagree with an interpretation that is nonetheless well grounded. And negative case analysis actively searches for data that contradict the emerging account and either revises it or explains the exception.

Transferability asks whether a reader can judge how far the findings apply elsewhere. It is not the researcher's claim but the reader's judgement, and it is enabled by thick description: enough detail about the setting, the participants, the period, and the conditions for someone in another context to assess similarity. A study that reports only themes without context has made transferability impossible to assess, which is a failure even though generalisation was never claimed.

Dependability asks whether the process was systematic and traceable. It is supported by an audit trail: a record of how the sample was built, how codes were developed and revised, what decisions were made and when, and why the analysis stopped where it did. In a thesis this appears as a documented coding procedure and an appendix showing the codebook's development.

Confirmability asks whether the findings are grounded in the data rather than in the researcher's preferences. It is supported by reflexivity, meaning explicit disclosure of the researcher's position, prior beliefs, and relationship to the setting, and by presenting raw data extracts so a reader can see the material from which an interpretation was drawn. Quotations in a qualitative findings chapter are evidence rather than illustration, and their function is exactly this.

Two warnings. Do not claim reliability in the statistical sense for qualitative work; the concept does not apply and using the word signals a misunderstanding. And do not list all four criteria and then describe no procedure for any of them, which is the most common form this section takes and the least convincing.

Bottom line: validity asks whether you measured the construct, reliability asks whether the measurement repeats, and reliable can be invalid while valid cannot be unreliable. Address content, construct, criterion, internal, and external validity under labelled headings. Report reliability per scale from your own data, and address low coefficients rather than omitting them. Name the specific threats your design faces, common method bias above all, and mitigate them in the design rather than in the limitations. In qualitative work use credibility, transferability, dependability, and confirmability, and pair each with the procedure that supports it.
الخامس

المعايير النوعية الأربعة

البحث النوعي لا يدّعي أن باحثًا آخر سيُنتج نتائج متطابقة، فيلزم مفردات مختلفة. والمعايير الأربعة مطالِبة لا متساهلة، وكلٌّ يسنده إجراء محدَّد.

المصداقيةقابلية النقلالاعتماديةإمكانية التأكيدهل النتائج مقنعة لمن فيالموقع؟هل يستطيع القارئ الحكمعلى انطباقها؟هل العملية قابلةللتتبع ومتسقة؟هل النتائج مجذّرة فيالبيانات لا في الباحث؟مراجعة المشاركينوالتثليثوصف كثيف للسياقمسار تدقيق للقراراتالانعكاسية ومقتطفاتخامالمكافئات النوعية؛ كلماتٌ مختلفة للهاجس الكامن نفسه

والمصداقية تسأل هل النتائج عرضٌ مقنع لواقع المشاركين. وثلاثة إجراءات تسندها. التثليث يقارن العروض عبر المصادر أو المناهج أو الباحثين، وقيمته ليست الاتفاق بل تفسير الاختلاف. ومراجعة المشاركين تعيد التفسير إلى المشاركين وتسأل هل يتعرفون فيه على خبرتهم، وهذا قوي ويجب التعامل معه بعناية، لأن المشاركين قد يخالفون تفسيرًا هو مع ذلك مجذّر جيدًا. وتحليل الحالة السالبة يبحث فعليًا عن بيانات تناقض العرض الناشئ ثم يراجعه أو يفسّر الاستثناء.

وقابلية النقل تسأل هل يستطيع القارئ الحكم إلى أي مدى تنطبق النتائج في مواضع أخرى. وهي ليست ادّعاء الباحث بل حكم القارئ، ويتيحها الوصف الكثيف: تفصيلٌ كافٍ عن الموقع والمشاركين والحقبة والشروط ليقيّم شخص في سياق آخر التشابه. والدراسة التي تُبلغ بالموضوعات وحدها بلا سياق جعلت تقييم قابلية النقل مستحيلًا، وهذا إخفاق وإن لم يُدَّعَ التعميم قط.

والاعتمادية تسأل هل كانت العملية منهجية وقابلة للتتبع. ويسندها مسار التدقيق: سجلٌ لكيفية بناء العيّنة، وكيف طُوِّرت الرموز ونُقِّحت، وأي قرارات اتُّخذت ومتى، ولماذا توقف التحليل حيث توقف. وفي أطروحة يظهر هذا إجراءَ ترميز موثَّقًا وملحقًا يُظهر تطور دليل الترميز.

وإمكانية التأكيد تسأل هل النتائج مجذّرة في البيانات لا في تفضيلات الباحث. وتسندها الانعكاسية، أي الإفصاح الصريح عن موقع الباحث ومعتقداته السابقة وعلاقته بالموقع، وعرضُ مقتطفات بيانات خام ليرى القارئ المادة التي استُخلص منها تفسير. والاقتباسات في فصل نتائج نوعي أدلةٌ لا توضيح، ووظيفتها هذه بالضبط.

وتحذيران. لا تدّعِ الثبات بالمعنى الإحصائي للعمل النوعي؛ فالمفهوم لا ينطبق واستعمال الكلمة يشير إلى سوء فهم. ولا تسرد المعايير الأربعة ثم لا تصف إجراءً لأي منها، وهذه أشيع صيغة يتخذها هذا القسم وأقلها إقناعًا.

الخلاصة: الصدق يسأل هل قِست البناء، والثبات يسأل هل يتكرر القياس، والثابت قد يكون غير صادق والصادق لا يمكن أن يكون غير ثابت. عالج صدق المحتوى والبناء والمحك والداخلي والخارجي تحت عناوين مُسمّاة. وأبلغ بالثبات لكل مقياس من بياناتك، وعالج المعاملات المنخفضة بدل إغفالها. وسمِّ التهديدات المحددة التي يواجهها تصميمك، وتحيّزَ المنهج المشترك قبل كل شيء، وخفّفها في التصميم لا في الحدود. وفي العمل النوعي استخدم المصداقية وقابلية النقل والاعتمادية وإمكانية التأكيد، وقرِن كلًّا بالإجراء الذي يسنده.