Every study faces the same pair of questions: who exactly will you study, and how many of them are enough? Sampling answers the first, and the logic of sample size, statistical power on one side, saturation on the other, answers the second. These are among the most examined decisions in any thesis, because they determine what the findings are allowed to claim, and they are also among the most misunderstood, because the two research traditions answer them with entirely different logics that are constantly confused with each other.
The quantitative logic is representation. The sample stands in for a population, so the central question is whether it mirrors that population well enough for conclusions to transfer, and the gold standard is random selection, which makes the mirror trustworthy by mathematics rather than hope. The qualitative logic is information. Cases are chosen because of what they can teach, not because they resemble an average, so the central question is whether the chosen cases are rich, relevant, and varied enough to answer the question deeply, and the stopping rule is not a number fixed in advance but saturation, the point where new cases stop teaching new things.
Confusing the two logics produces the classic examination failures: judging an interview study of twelve participants by survey standards, defending a convenience sample with the vocabulary of randomness, or announcing saturation without evidence. This framework sets out both families in depth with a decision diagram, treats sample size honestly in both traditions, gives saturation the careful treatment it rarely receives, and closes with a worked example and references. It is my own synthesis, written in my own words and grounded in recognized scholarship.
كل دراسة تواجه سؤالين متلازمين: من بالضبط ستدرس؟ وكم عددًا منهم يكفي؟ المعاينة تجيب عن السؤال الأول، ومنطق حجم العيّنة، القوة الإحصائية في جانب والتشبّع في الجانب الآخر، يجيب عن الثاني. وهذان القراران من أكثر ما يُناقَش في أي أطروحة، لأنهما يحدّدان ما يحق للنتائج أن تدّعيه، وهما أيضًا من أكثر المفاهيم التي يُساء فهمها، لأن تقليدَي البحث يجيبان عنهما بمنطقين مختلفين تمامًا يجري الخلط بينهما باستمرار.
المنطق الكمّي هو التمثيل. فالعيّنة تنوب عن مجتمع، والسؤال المركزي هو: هل تعكس العيّنة ذلك المجتمع بدقة تكفي لنقل الاستنتاجات إليه؟ والمعيار الذهبي هو الاختيار العشوائي، الذي يجعل هذا التمثيل موثوقًا بالرياضيات لا بالأمنيات. أما المنطق النوعي فهو المعلومة. فالحالات تُختار لما تستطيع أن تعلّمنا إياه، لا لأنها تشبه المتوسط، والسؤال المركزي هو: هل الحالات المختارة غنية وذات صلة ومتنوعة بما يكفي للإجابة عن السؤال بعمق؟ وقاعدة التوقف ليست رقمًا يُحدَّد مسبقًا بل التشبّع، وهو النقطة التي تتوقف عندها الحالات الجديدة عن تعليمنا شيئًا جديدًا.
والخلط بين المنطقين ينتج أخطاء المناقشات المشهورة: الحكم على دراسة مقابلات باثني عشر مشاركًا بمعايير المسوح، أو الدفاع عن عيّنة ميسّرة بمفردات العشوائية، أو إعلان التشبّع بلا دليل. يعرض هذا الإطار العائلتين بالتفصيل مع مخطط قرار، ويعالج حجم العيّنة بصدق في التقليدين، ويمنح التشبّع المعالجة الدقيقة التي نادرًا ما يحصل عليها، ويختم بمثال تطبيقي ومراجع. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
The diagram maps the sampling landscape: two families, each with its main members, each serving a different kind of claim. Knowing the members, and what each is for, turns sampling from jargon into a design decision.
In the probability family, every member of the population has a known, non-zero chance of selection, which is what licenses statistical inference. Simple random sampling draws names as from a hat and is the conceptual baseline. Stratified sampling first divides the population into groups that matter, departments, regions, grades, then samples randomly within each, guaranteeing that small but important groups are properly represented rather than left to chance. Cluster sampling randomly selects whole units, schools, branches, neighbourhoods, then studies people within them, trading some precision for enormous practicality when the population is scattered. The family's constant enemy is the sampling frame: the list from which selection happens. A perfect random draw from an incomplete list, missing new employees, or people without email, inherits every gap in the list, which is coverage error doing its quiet work.
In the purposive family, cases are chosen deliberately for what they can teach. Maximum-variation sampling picks cases as different as possible, so patterns that survive the variation are robust. Extreme-case sampling studies the outliers, the spectacular success, the repeated failure, because mechanisms show themselves at the edges. Typical-case sampling documents the ordinary run of a phenomenon. Snowball sampling asks each participant to open doors to the next, indispensable for reaching hidden or hard-to-access populations, at the known cost that the chain stays within connected circles. Theoretical sampling, grounded theory's own member, lets the developing analysis choose each next case. And convenience sampling, taking whoever is easiest to reach, is the family's weak member: sometimes unavoidable, never a virtue, and honest reporting names it and states the limits it imposes rather than dressing it in purposive language.
المخطط يرسم خريطة المعاينة: عائلتان، لكل واحدة أعضاؤها الرئيسيون، وكل واحدة تخدم نوعًا مختلفًا من الادّعاءات. ومعرفة هذه الأنواع، ووظيفة كل نوع، تحوّل المعاينة من مصطلحات جامدة إلى قرار تصميم واعٍ.
في العائلة الاحتمالية، لكل فرد في المجتمع فرصة اختيار معلومة وغير صفرية، وهذا هو ما يبيح الاستدلال الإحصائي. فالمعاينة العشوائية البسيطة تسحب الأسماء كما تُسحَب من قبعة، وهي الأساس النظري. والمعاينة الطبقية تقسم المجتمع أولًا إلى فئات مهمة، أقسام أو مناطق أو درجات وظيفية، ثم تختار عشوائيًا داخل كل فئة، فتضمن تمثيل الفئات الصغيرة المهمة تمثيلًا صحيحًا بدل تركها للصدفة. والمعاينة العنقودية تختار وحدات كاملة عشوائيًا، مدارس أو فروعًا أو أحياء، ثم تدرس الأفراد داخلها، فتضحّي بشيء من الدقة مقابل عملية أسهل بكثير عندما يكون المجتمع متناثرًا. والعدو الدائم لهذه العائلة هو إطار المعاينة: القائمة التي يجري الاختيار منها. فالسحب العشوائي المثالي من قائمة ناقصة، تفتقد الموظفين الجدد أو من لا بريد إلكتروني لهم، يرث كل فجوة في القائمة، وهذا هو خطأ التغطية وهو يعمل بصمت.
وفي العائلة القصدية، تُختار الحالات عمدًا لما تستطيع أن تعلّمنا إياه. فمعاينة أقصى تباين تختار حالات مختلفة قدر الإمكان، بحيث تكون الأنماط التي تصمد رغم هذا الاختلاف أنماطًا قوية. ومعاينة الحالات القصوى تدرس الحالات الاستثنائية، النجاح الباهر أو الفشل المتكرر، لأن الآليات تكشف نفسها عند الأطراف. ومعاينة الحالات النمطية توثّق المسار العادي للظاهرة. ومعاينة كرة الثلج تطلب من كل مشارك أن يفتح الباب للمشارك التالي، وهي أساسية للوصول إلى الفئات الخفية أو صعبة الوصول، مع ثمن معروف هو أن السلسلة تبقى داخل الدوائر المتعارفة. والمعاينة النظرية، وهي عضو النظرية المتجذّرة الخاص، تدع التحليل المتطور يختار الحالة التالية. أما معاينة الميسّر، أي أخذ من يسهل الوصول إليهم، فهي العضو الضعيف في العائلة: قد تكون أحيانًا لا مفرّ منها، لكنها ليست ميزة أبدًا، والتقرير الصادق يسمّيها باسمها ويذكر الحدود التي تفرضها بدل إلباسها لغة القصدية.
How many is enough has two honest answers, one per tradition, and both are better than the folk numbers that circulate in their place. The quantitative answer is calculated; the qualitative answer is evidenced.
In quantitative work the answer comes from power analysis, and its logic deserves to be understood rather than outsourced. Statistical power is the probability of detecting an effect that truly exists; it rises with sample size and with the size of the effect being hunted, and falls with noise. The researcher specifies the smallest effect worth caring about, the intended test, and the conventional error rates, and the calculation returns the sample needed. Two consequences follow. First, sample size is driven by the effect, not the population: detecting a subtle effect needs a large sample whether the population is one thousand or one million, which surprises those who assume big populations demand big samples. Second, an underpowered study is not merely weaker; it is systematically misleading, because the effects it does manage to detect are inflated by the selection of luck, the winner's curse of small samples. Reporting the power analysis, its assumptions, and the achieved sample against the target is the professional standard.
In qualitative work the honest answer is: until saturation, evidenced. Folk numbers circulate, twelve interviews, twenty, thirty, and reviews of actual studies show most themes appearing early, with diminishing novelty after roughly a dozen rich interviews in homogeneous groups; but these are observations about tendencies, not rules. The defensible practice is to plan a provisional range grounded in comparable studies, state the stopping criterion in advance as saturation at the level of themes or categories, and then document the approach to saturation as the study proceeds, which the next section details. Sample size in qualitative work is a claim to be supported, not a quota to be filled, and a paragraph that says what was planned, what was done, and how the stopping point was recognized outperforms any magic number.
لسؤال «كم يكفي؟» جوابان صادقان، واحد لكل تقليد، وكلاهما أفضل من الأرقام الشعبية المتداولة بدلًا منهما. الجواب الكمّي يُحسَب حسابًا؛ والجواب النوعي يُدعَم بالدليل.
في البحث الكمّي يأتي الجواب من تحليل القوة الإحصائية، ومنطقه يستحق أن يُفهَم لا أن يوكَل لغيرك. القوة الإحصائية هي احتمال اكتشاف أثر موجود فعلًا؛ وهي ترتفع مع حجم العيّنة ومع حجم الأثر المطلوب اكتشافه، وتنخفض مع الضجيج. يحدّد الباحث أصغر أثر يستحق الاهتمام، والاختبار المقصود، ومعدلات الخطأ المتعارف عليها، فيعيد الحساب حجم العيّنة المطلوب. وتترتب على ذلك نتيجتان. الأولى: حجم العيّنة يحدّده حجم الأثر لا حجم المجتمع؛ فاكتشاف أثر دقيق يحتاج عيّنة كبيرة سواء كان المجتمع ألف شخص أو مليونًا، وهذا يفاجئ من يظن أن المجتمعات الكبيرة تتطلب عيّنات كبيرة. والثانية: الدراسة ضعيفة القوة ليست أضعف فحسب، بل مضلّلة بشكل منهجي، لأن الآثار التي تنجح في اكتشافها تكون مضخّمة بفعل انتقاء الحظ، وهي ما تُعرَف بلعنة الفائز في العيّنات الصغيرة. والمعيار المهني هو ذكر تحليل القوة وافتراضاته والعيّنة المتحققة مقارنةً بالهدف.
وفي البحث النوعي الجواب الصادق هو: حتى التشبّع، مع الدليل عليه. تنتشر أرقام شعبية، اثنتا عشرة مقابلة أو عشرون أو ثلاثون، ومراجعات الدراسات الفعلية تُظهر أن معظم المواضيع تظهر مبكرًا، مع تناقص الجديد بعد نحو اثنتي عشرة مقابلة غنية في المجموعات المتجانسة؛ لكن هذه ملاحظات عن اتجاهات عامة، وليست قواعد. والممارسة السليمة هي التخطيط لنطاق مبدئي مبني على دراسات مشابهة، وذكر معيار التوقف مسبقًا بأنه التشبّع على مستوى المواضيع أو الفئات، ثم توثيق الاقتراب من التشبّع أثناء سير الدراسة، وهو ما يفصّله القسم التالي. فحجم العيّنة في البحث النوعي ادّعاء يجب دعمه، لا حصة يجب ملؤها، وفقرة تقول ما الذي خُطّط له وما الذي جرى وكيف عُرفت نقطة التوقف تتفوق على أي رقم سحري.
Saturation is the most invoked and least demonstrated concept in qualitative research. Used properly, it is a precise idea with observable evidence; used loosely, it is a permission slip for stopping when the researcher is tired. The difference is documentation.
The precise idea: saturation is a property of the analysis, not of the dataset. A category or theme is saturated when additional data yield no new properties, dimensions, or relationships for it, only repetition of what is already understood. Three consequences follow. Saturation is judged category by category, so a study can be saturated on its central themes while an interesting peripheral theme remains open, and honest reporting says so. Saturation depends on analysis keeping pace with collection, because a researcher who collects twenty interviews before coding any of them cannot have watched saturation approach; the concept presupposes the iterative loop. And saturation is relative to the question: the same data can saturate a broad descriptive question quickly while leaving a subtle process question hungry for more cases.
The observable evidence is what converts a claim into a demonstration. A saturation table tracks, wave by wave, how many new codes each round of interviews produced, and shows the curve flattening: eleven new codes from the first three interviews, five from the next three, one from the next, none from the last. Memos record the judgment as it formed: after interview nine, the resistance theme is stable; interviews ten and eleven added nothing to it. Negative-case hunting is reported: the final two participants were selected precisely because they seemed most likely to break the pattern, and did not. Where saturation was not reached on a theme, the limitation is stated plainly. This apparatus costs a page of the thesis, and it transforms the sample-size defence from an assertion an examiner can doubt into evidence an examiner can only read.
التشبّع هو المفهوم الأكثر استخدامًا والأقل إثباتًا في البحث النوعي. فإذا استُخدم كما ينبغي كان فكرة دقيقة لها دليل ملموس؛ وإذا استُخدم بتساهل صار رخصة للتوقف عندما يتعب الباحث. والفرق بين الحالتين هو التوثيق.
الفكرة الدقيقة: التشبّع صفة للتحليل لا لمجموعة البيانات. فالفئة أو الموضوع يتشبّع عندما لا تضيف البيانات الجديدة خصائص أو أبعادًا أو علاقات جديدة له، بل مجرد تكرار لما صار مفهومًا. وتترتب على ذلك ثلاث نتائج. الأولى: يُحكَم على التشبّع فئةً فئة، فقد تتشبّع الدراسة في مواضيعها المركزية بينما يبقى موضوع جانبي مثير مفتوحًا، والتقرير الصادق يقول ذلك صراحة. والثانية: التشبّع يعتمد على سير التحليل بالتوازي مع جمع البيانات، فالباحث الذي يجمع عشرين مقابلة قبل أن يرمّز أيًا منها لا يمكن أن يكون قد راقب اقتراب التشبّع؛ فالمفهوم يفترض الحلقة التكرارية. والثالثة: التشبّع نسبي بالنسبة للسؤال؛ فالبيانات نفسها قد تُشبع سؤالًا وصفيًا عامًا بسرعة بينما تترك سؤال عمليةٍ دقيقًا محتاجًا لحالات أكثر.
والدليل الملموس هو ما يحوّل الادّعاء إلى إثبات. جدول التشبّع يتتبع، دفعةً دفعة، كم رمزًا جديدًا أنتجته كل جولة مقابلات، ويُظهر المنحنى وهو يستوي: أحد عشر رمزًا جديدًا من المقابلات الثلاث الأولى، وخمسة من الثلاث التالية، وواحد من التي تليها، ولا شيء من الأخيرة. والمذكرات تسجّل الحكم وهو يتكوّن: بعد المقابلة التاسعة، استقر موضوع المقاومة؛ ولم تضف المقابلتان العاشرة والحادية عشرة شيئًا إليه. ويُذكَر البحث عن الحالات السالبة: اختير المشاركان الأخيران تحديدًا لأنهما بدوا الأكثر احتمالًا لكسر النمط، فلم يكسراه. وحيث لم يتحقق التشبّع في موضوع ما، يُذكَر هذا الحد بوضوح. هذا الجهاز كله يكلّف صفحة واحدة من الأطروحة، لكنه يحوّل الدفاع عن حجم العيّنة من ادّعاء يستطيع الممتحن التشكيك فيه إلى دليل لا يملك الممتحن إلا قراءته.
One mixed study, two sampling stories done properly. The project: measuring the spread of a new work practice across a large organization, then understanding how adopting teams actually live it.
The survey strand samples for representation. The population is defined as all 4,200 employees; the frame is the HR roster, checked for the known gaps of contractors and new joiners, both declared. Power analysis for the smallest effect of interest returns a target of 350 completed responses; anticipating half of invitees will respond, 700 are invited by stratified random sampling across the five divisions, guaranteeing the small research division its proportional voice. The achieved 384 responses are compared with the roster on division, tenure, and grade, one youthful skew is found, reported, and handled in analysis. Every sentence of that story is a defence an examiner cannot puncture, because each decision is named, reasoned, and checked.
The interview strand samples for information. From the survey's own results, adopting teams are listed and sixteen candidate teams identified; maximum variation selects six across division, team size, and adoption timing, with the stated reason that patterns surviving this variation will be robust. Within teams, snowball referrals recruit the informal influencers the org chart does not show. Analysis runs alongside collection; the saturation table shows new codes flattening after the fourth team, and teams five and six, chosen as most likely to break the pattern, one remote, one that adopted under protest, confirm rather than break it, with the protest team adding one new boundary condition, reported as such. The thesis states that the central process reached saturation while one peripheral theme, effects on client relationships, did not, and marks it for future work. Two strands, two logics, each defended in its own terms: that is sampling done properly, and it reads exactly as unglamorous and solid as good methodology should.
دراسة مختلطة واحدة، وقصّتا معاينة منفّذتان كما ينبغي. المشروع: قياس انتشار ممارسة عمل جديدة في منظمة كبيرة، ثم فهم كيف تعيشها الفرق المتبنّية فعليًا.
خيط المسح يعايِن للتمثيل. يُعرَّف المجتمع بأنه جميع الموظفين البالغ عددهم 4,200؛ والإطار هو سجل الموارد البشرية، بعد فحصه من الفجوات المعروفة وهي المتعاقدون والملتحقون الجدد، مع الإعلان عن كلتيهما. تحليل القوة لأصغر أثر مهم يعطي هدفًا هو 350 استجابة مكتملة؛ وتوقعًا لاستجابة نصف المدعوين، يُدعى 700 موظف بمعاينة عشوائية طبقية عبر الأقسام الخمسة، بما يضمن لقسم البحوث الصغير صوته النسبي. وتُقارَن الاستجابات المتحققة، وعددها 384، بالسجل من حيث القسم وسنوات الخدمة والدرجة، فيُكتشف ميل نحو الفئة الأصغر سنًا، ويُذكَر ويُعالَج في التحليل. كل جملة في هذه القصة دفاعٌ لا يستطيع الممتحن اختراقه، لأن كل قرار مسمًّى ومعلَّل ومفحوص.
وخيط المقابلات يعايِن للمعلومة. من نتائج المسح نفسها تُحصَر الفرق المتبنّية ويُحدَّد ستة عشر فريقًا مرشحًا؛ ويختار مبدأ أقصى تباين ستة فرق عبر القسم وحجم الفريق وتوقيت التبنّي، مع سبب معلن هو أن الأنماط التي تصمد رغم هذا التباين ستكون أنماطًا قوية. وداخل الفرق، تجنّد إحالات كرة الثلج المؤثرين غير الرسميين الذين لا يُظهرهم الهيكل التنظيمي. ويسير التحليل بالتوازي مع الجمع؛ ويُظهر جدول التشبّع الرموز الجديدة وهي تستوي بعد الفريق الرابع، بينما الفريقان الخامس والسادس، المختاران لأنهما الأكثر احتمالًا لكسر النمط، أحدهما عن بُعد والآخر تبنّى الممارسة معترضًا، يؤكدان النمط بدل كسره، مع إضافة فريق الاعتراض شرطًا حدّيًا واحدًا جديدًا يُذكَر بصفته تلك. وتذكر الأطروحة أن العملية المركزية بلغت التشبّع بينما لم يبلغه موضوع جانبي واحد، هو الأثر على علاقات العملاء، فيوضَع علامةً لبحث مستقبلي. خيطان، ومنطقان، وكل واحد مُدافَع عنه بمصطلحات تقليده: هكذا تكون المعاينة الصحيحة، وهي تُقرأ تمامًا كما ينبغي للمنهجية الجيدة أن تُقرأ: رصينة بلا بهرجة.