The experiment is the most powerful machine ever devised for answering one kind of question: does X cause Y? Everything else in the quantitative toolkit measures association; only experimental logic can, by its own structure, separate causation from the endless supply of rival explanations that haunt every observed correlation. Understanding that logic, and what to do when true experiments are impossible, which in social research is most of the time, is essential equipment for any doctoral researcher who intends to make, or even to evaluate, a causal claim.
The problem the experiment solves deserves a name: the counterfactual problem. To say the training caused the improvement is to claim that, had the same people not received the training, they would not have improved, but that counterfactual world is unobservable; the same person cannot both receive and not receive the treatment. Every causal design is an attempt to construct a credible substitute for the unobservable counterfactual, and designs differ precisely in how credible their substitute is. The true experiment's substitute, a randomly assigned control group, is the most credible ever invented, which is why the randomized controlled trial sits at the top of every evidence hierarchy.
But randomization is often impossible, unethical, or absurd in real organizational and social settings: one cannot randomly assign firms to bankruptcy, children to family backgrounds, or countries to policies. The quasi-experimental tradition, built by Campbell and colleagues, is the disciplined art of approaching causal inference without random assignment, and its designs, and the validity-threat vocabulary that accompanies them, are among the most practically valuable tools social science possesses. This framework sets out the causal logic, the true experimental designs, the major quasi-designs, the full catalogue of validity threats, and a worked example, closing with standards and references. It is my own synthesis, written in my own words and grounded in recognized scholarship.
التجربة أقوى آلةٍ ابتُكرت للإجابة عن نوعٍ واحد من الأسئلة: هل يسبّب X النتيجة Y؟ كل ما عداها في العدّة الكمّية يقيس الارتباط؛ والمنطق التجريبي وحده يستطيع، ببنيته ذاتها، فصل السببية عن المعين الذي لا ينضب من التفسيرات المنافسة التي تطارد كل ارتباطٍ ملحوظ. وفهم ذلك المنطق، وما يُفعَل حين تستحيل التجارب الحقيقية، وهو في البحث الاجتماعي معظم الوقت، عتادٌ ضروري لأي باحث دكتوراه ينوي إطلاق ادّعاءٍ سببي، أو حتى تقييمه.
والمشكلة التي تحلّها التجربة تستحقّ اسمًا: مشكلة المضادّ الواقعي. فقول «التدريب سبّب التحسّن» ادّعاءٌ أنه، لو لم يتلقَّ الأشخاص ذاتهم التدريب، لما تحسّنوا، لكن ذلك العالم المضادّ غير قابل للملاحظة؛ فالشخص ذاته لا يستطيع تلقّي المعالجة وعدم تلقّيها معًا. وكل تصميمٍ سببي محاولةٌ لبناء بديلٍ ذي مصداقية للمضادّ الواقعي غير الملحوظ، وتختلف التصاميم بالضبط في مصداقية بديلها. وبديل التجربة الحقيقية، مجموعة ضبطٍ موزَّعة عشوائيًا، الأعلى مصداقيةً على الإطلاق، ولهذا تجلس التجربة المضبوطة المعشّاة على قمّة كل هرم دليل.
لكن التعشية غالبًا مستحيلة أو غير أخلاقية أو عبثية في البيئات التنظيمية والاجتماعية الحقيقية: لا يمكن توزيع الشركات عشوائيًا على الإفلاس، ولا الأطفال على الخلفيات الأسرية، ولا الدول على السياسات. والتقليد شبه التجريبي، الذي بناه كامبل وزملاؤه، الفنّ المنضبط لمقاربة الاستدلال السببي دون توزيعٍ عشوائي، وتصاميمه، ومفردات تهديدات الصدق المرافقة لها، من أثمن أدوات العلم الاجتماعي عمليًا. يعرض هذا الإطار المنطق السببي، والتصاميم التجريبية الحقيقية، والتصاميم الشبيهة الكبرى، وفهرس تهديدات الصدق الكامل، ومثالًا مشتغَلًا، خاتمًا بالمعايير والمراجع. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
The true experiment's power comes from one move, random assignment, and understanding exactly what that move accomplishes explains both why experiments are trusted and what every other design is struggling to imitate. The diagram shows the machinery.
Suppose training volunteers outperform non-volunteers. The obvious rival explanation is selection: volunteers differ from non-volunteers before any training, in motivation, ability, ambition, and any of those pre-existing differences could produce the outcome gap by itself. This is the generic curse of comparing self-selected groups, and no amount of statistical adjustment can fully lift it, because adjustment handles only the differences one has measured, while the dangerous differences are the unmeasured ones. Random assignment dissolves the curse in one stroke: when chance alone decides who gets the treatment, the two groups are, in expectation, identical in every respect, measured and unmeasured, known and unknown, except the treatment itself. Any outcome difference beyond chance variation then has only one place to come from.
Three companions complete the design. The control group supplies the counterfactual: it shows what would have happened to the treated, had they not been treated, because it is, statistically, the same collection of people. Manipulation, the researcher actively administering the treatment rather than observing its natural occurrence, guarantees the causal arrow's direction, since an administered cause cannot have been produced by its own effect. And controlled conditions, holding everything else as constant as the setting allows, reduce noise so the effect can be seen. Where all three are present with randomization, causal inference is as strong as empirical research gets; every design in the rest of this page is a strategy for living without one or more of them.
قوّة التجربة الحقيقية تأتي من حركةٍ واحدة، التوزيع العشوائي، وفهم ما تنجزه تلك الحركة بالضبط يفسّر لماذا تُوثَق التجارب وما الذي يجاهد كل تصميمٍ آخر لتقليده. والمخطط يُظهِر الآلة.
افترض أن متطوّعي التدريب يتفوّقون على غير المتطوّعين. التفسير المنافس البديهي الانتقاء: المتطوّعون يختلفون عن غيرهم قبل أي تدريب، في الدافعية والقدرة والطموح، وأيٌّ من تلك الفروق المسبقة يستطيع وحده إنتاج فجوة النتيجة. هذه لعنة مقارنة المجموعات ذاتية الانتقاء العامة، ولا قدر من التعديل الإحصائي يرفعها كاملةً، لأن التعديل يعالج الفروق المقيسة فقط، بينما الفروق الخطرة هي غير المقيسة. والتوزيع العشوائي يحلّ اللعنة بضربةٍ واحدة: حين تقرّر الصدفة وحدها من يتلقّى المعالجة، تكون المجموعتان، بالتوقّع، متطابقتين في كل جانب، مقيسًا وغير مقيس، معلومًا ومجهولًا، عدا المعالجة ذاتها. وأي فرق نتيجةٍ يتجاوز تباين الصدفة لا يبقى له إلّا مصدرٌ واحد.
وثلاثة رفاق يكملون التصميم. مجموعة الضبط تورّد المضادّ الواقعي: تُظهِر ما كان سيحدث للمعالَجين لو لم يُعالَجوا، لأنها، إحصائيًا, المجموعة ذاتها من الناس. والمعالجة الفاعلة، إدارة الباحث المعالجة بنشاطٍ بدل ملاحظة وقوعها الطبيعي، تضمن اتجاه السهم السببي، إذ السبب المُدار لا يمكن أن يكون نتاج أثره. والظروف المضبوطة، تثبيت كل شيءٍ آخر بقدر ما تسمح البيئة، تخفض الضجيج ليُرى الأثر. وحيث تحضر الثلاثة مع التعشية، يكون الاستدلال السببي أقوى ما يبلغه البحث التجريبي؛ وكل تصميمٍ في بقية هذه الصفحة استراتيجية عيشٍ دون واحدٍ منها أو أكثر.
When random assignment is out of reach, the quasi-experimental repertoire offers designs of graded strength, each buying back some of the lost inferential power through structure. Knowing the repertoire converts impossible cannot randomize situations into designable studies.
The weakest members are cautionary tales. The one-group post-test design, treat, then measure, supports almost no inference, since nothing shows what the group looked like before or would have looked like without treatment. Adding a pre-test, the one-group pre-post design, shows change but cannot attribute it: anything else happening at the same time could be the real cause. The workhorse of practice is the non-equivalent control group design: a treated group and an untreated comparison group, both measured before and after, where the comparison group was not randomly formed, an intact department, a neighbouring school, a matched firm. Its strength depends entirely on how comparable the comparison is, and the pre-test is its safety check, revealing how far apart the groups started.
Three stronger designs deserve their reputations. The interrupted time-series design measures the outcome many times before and after the intervention, so the intervention must visibly break an established trend; a policy that bends a ten-year curve at exactly the month of adoption is hard to explain away. The regression discontinuity design exploits a cutoff rule: when a programme is assigned by threshold, scholarship above a score, aid below an income line, cases just either side of the cutoff are effectively identical except for treatment, creating a local randomized experiment at the boundary. And difference-in-differences compares the change in a treated group with the change in an untreated group over the same period, subtracting out shared trends: what grew in the treated beyond what grew in the comparison is the effect estimate, provided the two would otherwise have moved in parallel, an assumption to argue, not assume. Natural experiments, where policy accidents or external shocks assign treatment almost as if by chance, extend the family further and reward the researcher who recognizes one in the wild.
حين يتعذّر التوزيع العشوائي، تقدّم ذخيرة شبه التجريبي تصاميم بقوًى متدرّجة، كلٌّ يستردّ بعض القوّة الاستدلالية المفقودة عبر البنية. ومعرفة الذخيرة تحوّل مواقف «لا أستطيع التعشية» المستحيلة إلى دراساتٍ قابلة للتصميم.
الأعضاء الأضعف عِبرٌ تحذيرية. تصميم المجموعة الواحدة بقياسٍ بعدي، عالِج ثم قِس، لا يدعم استدلالًا تقريبًا، إذ لا شيء يُظهِر كيف بدت المجموعة قبلًا أو كيف كانت ستبدو دون معالجة. وإضافة قياسٍ قبلي، تصميم قبلي-بعدي لمجموعةٍ واحدة، تُظهِر تغيّرًا لكنها لا تستطيع عزوه: فأي شيءٍ آخر يحدث في الوقت ذاته قد يكون السبب الحقيقي. وحصان عمل الممارسة تصميم مجموعة الضبط غير المكافئة: مجموعةٌ معالَجة وأخرى مقارنة غير معالَجة، كلتاهما مقيسة قبلًا وبعدًا، حيث لم تُشكَّل المقارنة عشوائيًا، قسمٌ قائم، مدرسةٌ مجاورة، شركةٌ مطابَقة. وقوّته تعتمد كليًا على قابلية المقارنة للمقارنة، والقياس القبلي صمّام أمانه، كاشفًا كم تباعدت المجموعتان بدايةً.
وثلاثة تصاميم أقوى تستحقّ سمعتها. تصميم السلاسل الزمنية المقطوعة يقيس النتيجة مراتٍ كثيرة قبل التدخّل وبعده، فيجب أن يكسر التدخّل اتجاهًا راسخًا كسرًا مرئيًا؛ وسياسةٌ تثني منحنى عشر سنواتٍ عند شهر التبنّي بالضبط يصعب تفسيرها بعيدًا. وتصميم انقطاع الانحدار يستغلّ قاعدة قطع: حين يُوزَّع برنامجٌ بعتبة، منحةٌ فوق درجة، إعانةٌ تحت خطّ دخل، تكون الحالات على جانبَي القطع مباشرةً متطابقةً فعليًا عدا المعالجة، خالقةً تجربةً معشّاة محلية عند الحدّ. وفروق الفروق تقارن تغيّر مجموعةٍ معالَجة بتغيّر أخرى غير معالَجة في الفترة ذاتها، طارحةً الاتجاهات المشتركة: ما نما في المعالَجة فوق ما نما في المقارنة هو تقدير الأثر، شرط أن الاثنتين كانتا ستتحرّكان متوازيتين لولاه، افتراضٌ يُحاجّ لا يُفترَض. والتجارب الطبيعية، حيث توزّع صدف السياسات أو الصدمات الخارجية المعالجة شبه صدفةٍ، تمدّ العائلة أبعد وتكافئ الباحث الذي يتعرّف على واحدةٍ في البرّية.
Campbell's great methodological gift was a shared vocabulary of validity threats: named rival explanations that any causal claim must survive. The catalogue works as a checklist, and running a design against it before data collection is the cheapest quality control research offers.
The internal validity threats attack the causal claim itself. Selection: the groups differed before treatment, the master threat of all non-randomized comparison. History: an outside event, coinciding with the treatment, produced the change. Maturation: the participants were changing anyway, growing, learning, tiring, and time is the real cause. Testing: the pre-test itself changed behaviour. Instrumentation: the measure drifted between waves, so apparent change is measurement change. Regression to the mean: groups selected for extreme scores drift back toward average by pure statistics, mimicking an effect, the classic trap of studying the worst performers and celebrating their inevitable improvement. Attrition: dropouts differ from stayers, silently rebuilding the selection problem inside a randomized design. Each named threat is a question to ask of any study, including one's own, and the discipline is to address each explicitly rather than hope the reader does not think of it.
| Threatالتهديد | The rival storyالقصّة المنافسة |
|---|---|
| Selectionالانتقاء | groups differed before treatmentاختلفت المجموعات قبل المعالجة |
| Historyالتاريخ | a coinciding outside event caused itحدثٌ خارجي متزامن سبّبها |
| Maturationالنضج | they were changing anywayكانوا يتغيّرون على أي حال |
| Regressionالانحدار للمتوسط | extremes drift back by statistics aloneالقيم القصوى ترتدّ إحصائيًا وحدها |
| Attritionالتسرّب | dropouts rebuilt the selection problemالمتسرّبون أعادوا بناء مشكلة الانتقاء |
External validity asks the further question: granted the effect is real here, for whom, where, and when does it hold? Laboratory purity, volunteer samples, and single-site studies all narrow the claim's reach, and the trade-off is structural: the controls that strengthen internal validity often make the setting less like the world the claim is meant to serve. Construct validity asks whether the operations really embody the concepts, is the treatment actually the theoretical variable, or something bundled with it, placebo effects and experimenter expectancy being the classic bundlings. A complete causal argument addresses all three families, in that order of priority, because an effect must first be real before its reach and meaning are worth debating.
هدية كامبل المنهجية الكبرى مفردات تهديدات صدقٍ مشتركة: تفسيراتٌ منافسة مسمّاة يجب أن ينجو منها أي ادّعاءٍ سببي. ويعمل الفهرس قائمة فحص، وإجراء التصميم عليه قبل جمع البيانات أرخص ضبط جودةٍ يقدّمه البحث.
تهديدات الصدق الداخلي تهاجم الادّعاء السببي ذاته. الانتقاء: اختلفت المجموعات قبل المعالجة، التهديد الأكبر لكل مقارنةٍ غير معشّاة. والتاريخ: حدثٌ خارجي، تزامن مع المعالجة، أنتج التغيّر. والنضج: كان المشاركون يتغيّرون أصلًا، ينمون ويتعلّمون ويتعبون، والزمن السبب الحقيقي. والاختبار: القياس القبلي نفسه غيّر السلوك. والأدوات: انجرف المقياس بين الموجات، فالتغيّر الظاهر تغيّر قياس. والانحدار للمتوسط: المجموعات المختارة لدرجاتٍ قصوى ترتدّ نحو المتوسط بالإحصاء الخالص، محاكيةً أثرًا، الفخّ الكلاسيكي لدراسة أسوأ المؤدّين والاحتفاء بتحسّنهم الحتمي. والتسرّب: المنسحبون يختلفون عن الباقين، معيدين بصمتٍ بناء مشكلة الانتقاء داخل تصميمٍ معشًّى. كل تهديدٍ مسمًّى سؤالٌ يُطرَح على أي دراسة، بما فيها دراسة المرء، والانضباط معالجة كلٍّ صراحةً لا الأمل ألّا يفكّر فيه القارئ.
| Threatالتهديد | The rival storyالقصّة المنافسة |
|---|---|
| Selectionالانتقاء | groups differed before treatmentاختلفت المجموعات قبل المعالجة |
| Historyالتاريخ | a coinciding outside event caused itحدثٌ خارجي متزامن سبّبها |
| Maturationالنضج | they were changing anywayكانوا يتغيّرون على أي حال |
| Regressionالانحدار للمتوسط | extremes drift back by statistics aloneالقيم القصوى ترتدّ إحصائيًا وحدها |
| Attritionالتسرّب | dropouts rebuilt the selection problemالمتسرّبون أعادوا بناء مشكلة الانتقاء |
والصدق الخارجي يسأل السؤال الإضافي: بافتراض أن الأثر حقيقي هنا، لمن وأين ومتى يصمد؟ نقاء المختبر وعيّنات المتطوّعين ودراسات الموقع الواحد كلها تضيّق مدى الادّعاء، والمبادلة بنيوية: فالضوابط التي تقوّي الصدق الداخلي كثيرًا ما تجعل البيئة أقلّ شبهًا بالعالم الذي يُراد للادّعاء خدمته. وصدق البناء يسأل أتجسّد العملياتُ المفاهيمَ فعلًا، أالمعالجة هي المتغيّر النظري حقًا أم شيءٌ محزوم معه، وآثار الوهم وتوقّعات المجرّب الحزم الكلاسيكية. والحجّة السببية الكاملة تعالج العائلات الثلاث، بذلك الترتيب من الأولوية، لأن الأثر يجب أولًا أن يكون حقيقيًا قبل أن يستحقّ مداه ومعناه النقاش.
Follow one evaluation through three successive designs and watch the causal claim strengthen as structure is added. The intervention: a company rolls out a mentoring programme intended to reduce first-year employee turnover.
Design one, what the company actually did first: offer mentoring to volunteers, then compare turnover between mentored and unmentored employees. Mentored turnover is half the unmentored rate, and the HR report declares success. The catalogue dismantles the claim in seconds: selection is fatal, because the kind of newcomer who volunteers for mentoring, engaged, ambitious, already committed, is exactly the kind who stays anyway. History and maturation lurk behind it. This design cannot distinguish the programme's effect from the volunteers' character, and no statistical patching of measured covariates rescues it, because commitment itself was never measured.
Design two, the feasible quasi-experiment: the programme rolls out site by site for operational reasons. Two comparable sites are found, one early-adopting and one not yet started, with three years of pre-programme turnover data showing parallel trends. The difference-in-differences logic now works: turnover fell in the adopting site by five points beyond the change in the comparison site over the same period. Selection at the individual level is bypassed because whole sites, not volunteers, are compared; history is partially handled because shared shocks hit both sites; the remaining vulnerability, a site-specific event coinciding with adoption, is checked by interviewing both sites about the period and by the time-series showing the break lands on the adoption month. Design three, the true experiment the evidence then justified: the next phase randomizes which new hires are offered mentoring within each site, with turnover tracked for a year. The randomized estimate, four points, close to the quasi-estimate, completes the chain, and the two designs together make a case far stronger than either alone: the experiment nails the internal validity, the quasi-design shows the effect survives in ordinary operational conditions.
تتبّع تقييمًا واحدًا عبر ثلاثة تصاميم متتابعة وراقب الادّعاء السببي يقوى مع إضافة البنية. التدخّل: تطلق شركةٌ برنامج إرشادٍ يُراد به خفض دوران موظفي السنة الأولى.
التصميم الأول، ما فعلته الشركة فعلًا أولًا: عرض الإرشاد على متطوّعين، ثم مقارنة الدوران بين المرشَدين وغير المرشَدين. دوران المرشَدين نصف معدّل غيرهم، ويعلن تقرير الموارد البشرية النجاح. والفهرس يفكّك الادّعاء في ثوانٍ: الانتقاء قاتل، لأن نوع الوافد الذي يتطوّع للإرشاد، مندمجًا طموحًا ملتزمًا سلفًا، هو بالضبط النوع الذي يبقى على أي حال. والتاريخ والنضج يتربّصان خلفه. هذا التصميم لا يستطيع تمييز أثر البرنامج عن شخصية المتطوّعين، ولا ترقيع إحصائي بمتغيّراتٍ مقيسة ينقذه، لأن الالتزام نفسه لم يُقَس قط.
التصميم الثاني، شبه التجريبي الممكن: يُطلَق البرنامج موقعًا موقعًا لأسبابٍ تشغيلية. يُوجَد موقعان قابلان للمقارنة، أحدهما مبكّر التبنّي والآخر لم يبدأ، مع ثلاث سنوات بيانات دورانٍ قبل البرنامج تُظهِر اتجاهاتٍ متوازية. ومنطق فروق الفروق يعمل الآن: هبط الدوران في الموقع المتبنّي خمس نقاطٍ فوق تغيّر موقع المقارنة في الفترة ذاتها. يُتجاوَز الانتقاء الفردي لأن المقارنة بين مواقع كاملةٍ لا متطوّعين؛ ويُعالَج التاريخ جزئيًا لأن الصدمات المشتركة تضرب الموقعين؛ والهشاشة الباقية، حدثٌ خاص بالموقع تزامن مع التبنّي، تُفحَص بمقابلة الموقعين عن الفترة وبالسلسلة الزمنية المُظهِرة أن الكسر يقع على شهر التبنّي. التصميم الثالث، التجربة الحقيقية التي برّرها الدليل بعدها: تعشّي المرحلة التالية أي الموظفين الجدد يُعرَض عليهم الإرشاد داخل كل موقع، بتتبّع الدوران سنة. والتقدير المعشّى، أربع نقاط، قريبٌ من الشبيه، يكمل السلسلة، والتصميمان معًا يبنيان قضيةً أقوى بكثيرٍ من أيٍّ وحده: التجربة تسمّر الصدق الداخلي، والشبيه يُظهِر أن الأثر ينجو في الظروف التشغيلية العادية.