Inferential statistics answer one question: how confident can you be that a pattern in your sample reflects something in the population rather than the accident of who happened to be sampled? Everything technical follows from that question, and most misuse follows from forgetting it.
The logic is indirect and worth stating precisely, because the standard shorthand hides it. You assume there is no effect in the population, which is the null hypothesis. You then compute how likely a result as extreme as yours would be if that assumption were true. If it is sufficiently unlikely, you conclude the assumption is implausible and reject it. You never prove that an effect exists; you establish that its absence is an uncomfortable explanation for what you observed.
That indirectness generates the two errors most commonly made with p values. A p value is not the probability that the null hypothesis is true, and it is not a measure of how large or important the effect is. It is the probability of data at least this extreme under an assumption. A very small p value in a very large sample can accompany an effect too small to matter, which is why effect sizes are reported alongside and why a result that is statistically significant may be practically irrelevant.
This framework covers the logic of hypothesis testing and the two error types, how to choose the right test, effect size and confidence intervals, the main techniques in business research, and how to report results without overclaiming. It is my own synthesis, written in my own words and grounded in recognized scholarship.
الإحصاء الاستدلالي يجيب سؤالًا واحدًا: كم يمكنك أن تثق بأن نمطًا في عيّنتك يعكس شيئًا في المجتمع لا مصادفةَ من صادف أن أُخذت منهم العيّنة؟ وكل ما هو تقني يتبع من ذلك السؤال، ومعظم سوء الاستعمال يتبع من نسيانه.
والمنطق غير مباشر ويستحق الذكر بدقة، لأن الاختصار المعياري يخفيه. فأنت تفترض ألا أثر في المجتمع، وهذه الفرضية الصفرية. ثم تحسب كم يُحتمَل أن تظهر نتيجة بتطرف نتيجتك لو صحّ ذلك الافتراض. فإن كان الاحتمال ضئيلًا بما يكفي، تستنتج أن الافتراض غير معقول وترفضه. ولا تبرهن أبدًا على وجود أثر؛ بل تثبت أن غيابه تفسيرٌ غير مريح لما لاحظت.
وتلك اللامباشرة تولّد الخطأين الأشيع مع القيم الاحتمالية. فالقيمة الاحتمالية ليست احتمالَ أن تكون الفرضية الصفرية صحيحة، وليست مقياسًا لكم الأثر كبير أو مهم. إنها احتمال بيانات بهذا التطرف على الأقل تحت افتراض. وقيمةٌ احتمالية صغيرة جدًا في عيّنة كبيرة جدًا قد يرافقها أثر أصغر من أن يهم، ولهذا يُبلَّغ بأحجام الأثر إلى جانبها ولهذا قد تكون نتيجةٌ دالة إحصائيًا غيرَ ذات صلة عمليًا.
ويغطي هذا الإطار منطق اختبار الفرضيات ونوعي الخطأ، وكيف يُختار الاختبار الصحيح، وحجم الأثر وفترات الثقة، والتقنيات الرئيسة في بحوث الأعمال، وكيف يُبلَّغ بالنتائج بلا ادّعاء زائد. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
Four steps, always the same, whatever the technique. Understanding them is what lets you interpret any test you have not met before.
The null hypothesis states that there is no effect: no difference between the groups, no relationship between the variables, no association between the categories. The alternative hypothesis states that there is one. Note that these are statements about the population, not about your sample; your sample always shows some difference, and the question is whether it is larger than sampling variation would produce.
The significance level, conventionally 0.05, is the threshold you fix in advance for how unlikely a result must be before you reject the null. It is a convention rather than a law of nature, and it encodes a tolerance: at 0.05, one test in twenty will produce a significant result when there is no effect at all. That fact has a consequence people often miss. Running twenty tests on data with no real effects produces, on average, one significant finding, and reporting that finding without mentioning the other nineteen is how spurious results enter the literature.
Two errors are possible and they trade off. A Type I error is rejecting a true null, meaning claiming an effect that does not exist, and its rate is the significance level you chose. A Type II error is failing to reject a false null, meaning missing an effect that does exist, and its rate depends on the sample size, the effect size, and the significance level. Power, conventionally set at 0.80, is one minus the Type II error rate: an eighty percent chance of detecting an effect of the size you expect, if it is there.
Three consequences follow for interpretation. Failing to reject the null is not evidence of no effect. It means your data were consistent with no effect, which in a small sample they would be even if a substantial effect existed. The honest phrasing is that no significant relationship was found, not that no relationship exists.
Significance depends on sample size. The same correlation of 0.15 is non-significant at n equals 100 and significant at n equals 500. Nothing about the world changed; only the precision of the estimate did. And one-tailed tests, which test for an effect in one direction only, are legitimate when the hypothesis is genuinely directional and stated before the analysis, and are a way of obtaining significance cheaply when they are not.
أربع خطوات، هي نفسها دائمًا، أيًّا كانت التقنية. وفهمها هو ما يتيح لك تفسير أي اختبار لم تصادفه من قبل.
والفرضية الصفرية تذكر ألا أثر: لا فرق بين المجموعات، ولا علاقة بين المتغيرات، ولا ارتباط بين الفئات. والفرضية البديلة تذكر أن ثمة واحدًا. ولاحظ أن هاتين عبارتان عن المجتمع لا عن عيّنتك؛ فعيّنتك تُظهر فرقًا ما دائمًا، والسؤال هل هو أكبر مما ينتجه تباين المعاينة.
ومستوى الدلالة، وهو 0.05 عرفًا، عتبةٌ تثبّتها سلفًا لكم يجب أن تكون نتيجة غير محتملة قبل رفض الصفرية. وهو عُرف لا قانون طبيعة، ويرمّز تسامحًا: فعند 0.05، سيُنتج اختبارٌ من كل عشرين نتيجةً دالة حين لا أثر البتة. ولتلك الواقعة نتيجة يُغفلها الناس كثيرًا. فتشغيل عشرين اختبارًا على بيانات بلا آثار حقيقية يُنتج، وسطيًا، نتيجةً دالة واحدة، والإبلاغ بتلك النتيجة دون ذكر التسعة عشر الأخرى هو كيف تدخل النتائج الزائفة الأدبيات.
وخطآن ممكنان ويتقايضان. فـخطأ النوع الأول رفضُ صفرية صحيحة، أي ادّعاءُ أثر غير موجود، ومعدله مستوى الدلالة الذي اخترت. وخطأ النوع الثاني الإخفاقُ في رفض صفرية خاطئة، أي تفويتُ أثر موجود، ومعدله يعتمد على حجم العيّنة وحجم الأثر ومستوى الدلالة. والقوة، المضبوطة عرفًا على 0.80، هي واحد ناقص معدل خطأ النوع الثاني: أي فرصة ثمانين بالمئة لكشف أثر بالحجم الذي تتوقعه، إن كان موجودًا.
وثلاث نتائج تتبع للتفسير. الإخفاق في رفض الصفرية ليس دليلًا على انعدام الأثر. بل يعني أن بياناتك كانت متسقة مع انعدام الأثر، وهذا ما ستكونه في عيّنة صغيرة حتى لو وُجد أثر كبير. والصياغة الصادقة أنه «لم يُوجَد ارتباط دال»، لا أنه «لا توجد علاقة».
والدلالة تعتمد على حجم العيّنة. فالارتباط نفسه البالغ 0.15 غير دال عند 100 ودال عند 500. ولم يتغير شيء في العالم؛ بل تغيرت دقة التقدير فقط. والاختبارات أحادية الذيل، التي تختبر أثرًا في اتجاه واحد فقط، مشروعة حين تكون الفرضية اتجاهية فعلًا ومذكورة قبل التحليل، وهي طريقة لنيل الدلالة برخص حين لا تكون كذلك.
Three questions choose the test: what kind of question you are asking, what level your variables are measured at, and whether the parametric assumptions hold.
Comparing two groups on a continuous outcome uses an independent samples t test when the groups are separate people and a paired t test when the same people are measured twice. Where the outcome is ordinal or the distribution is badly skewed in small samples, the rank-based alternative is appropriate and should be named rather than described as a non-parametric test.
Comparing three or more groups uses analysis of variance, which tests whether any of the group means differ. A significant result tells you that at least one pair differs and not which, which is why a post hoc comparison follows and must be reported. Running multiple t tests instead inflates the Type I error rate, which is precisely what the analysis of variance exists to control.
Two categorical variables use the chi-square test of independence, which asks whether the distribution across one variable's categories depends on the other's. Its main requirement is adequate expected cell counts, and where cells are sparse the categories should be combined or an exact test used.
Two continuous variables use correlation: Pearson for interval data with a roughly linear relationship, Spearman for ordinal data or monotonic but non-linear relationships. Correlation reports association and never direction, and the discipline of writing is associated with rather than affects belongs here more than anywhere.
Multiple regression is the workhorse of quantitative business research. It estimates the relationship between each predictor and the outcome while holding the others constant, which is what allows a claim that an association is not explained by the control variables included. Its output is a coefficient per predictor, its significance, and an overall measure of how much variance the model explains. Logistic regression does the same when the outcome is binary, and its coefficients are interpreted as changes in odds rather than in the outcome itself.
Two rules govern all of this. The question chooses the test, not the reverse. Write each hypothesis, then name the test that addresses it, in a table with one row per hypothesis. And check the assumptions before running it, since a technique applied to data that violate its requirements produces output that looks identical to valid output.
ثلاثة أسئلة تختار الاختبار: أي نوع من الأسئلة تسأل، وبأي مستوى قيست متغيراتك، وهل تصح الافتراضات المعلمية.
ومقارنة مجموعتين على مخرَج متصل تستخدم اختبار t لعيّنتين مستقلتين حين تكون المجموعتان أشخاصًا منفصلين واختبار t المزدوج حين يُقاس الأشخاص أنفسهم مرتين. وحيث يكون المخرَج رتبيًا أو التوزيع ملتويًا بشدة في عيّنات صغيرة، يلائم البديل الرتبي وينبغي تسميته لا وصفه بـ«اختبار لامعلمي».
ومقارنة ثلاث مجموعات فأكثر تستخدم تحليل التباين، الذي يختبر هل تختلف أي من متوسطات المجموعات. والنتيجة الدالة تخبرك أن زوجًا واحدًا على الأقل يختلف ولا تخبرك أيّها، ولهذا تتبعها مقارنة بعدية ويجب الإبلاغ بها. وتشغيل اختبارات t متعددة بدلًا من ذلك يضخّم معدل خطأ النوع الأول، وهذا بالضبط ما وُجد تحليل التباين لضبطه.
ومتغيران فئويان يستخدمان اختبار مربع كاي للاستقلال، الذي يسأل هل يعتمد التوزيع عبر فئات متغير على الآخر. ومتطلبه الرئيس أعدادُ خانات متوقعة كافية، وحيث تكون الخانات نادرة ينبغي دمج الفئات أو استخدام اختبار دقيق.
ومتغيران متصلان يستخدمان الارتباط: بيرسون للبيانات الفترية بعلاقة خطية تقريبًا، وسبيرمان للبيانات الرتبية أو للعلاقات الرتيبة غير الخطية. والارتباط يُبلغ بالاقتران ولا يُبلغ بالاتجاه أبدًا، وانضباطُ الكتابة «يرتبط بـ» لا «يؤثر في» ينتمي هنا أكثر من أي موضع.
والانحدار المتعدد حصان عمل بحوث الأعمال الكمية. فهو يقدّر العلاقة بين كل متنبئ والمخرَج مع تثبيت الآخرين، وهذا ما يتيح ادّعاء أن اقترانًا لا تفسّره متغيرات الضبط المدرَجة. ومخرَجه معاملٌ لكل متنبئ ودلالته ومقياسٌ كلي لكم يفسّر النموذج من التباين. والانحدار اللوجستي يفعل الشيء نفسه حين يكون المخرَج ثنائيًا، وتُفسَّر معاملاته تغيراتٍ في الأرجحية لا في المخرَج نفسه.
وقاعدتان تحكمان هذا كله. السؤال يختار الاختبار لا العكس. اكتب كل فرضية، ثم سمِّ الاختبار الذي يعالجها، في جدول بصف لكل فرضية. وافحص الافتراضات قبل تشغيله، لأن تقنيةً مطبَّقة على بيانات تنتهك متطلباتها تُنتج مخرَجًا يبدو مطابقًا للمخرَج الصحيح.
The p value answers whether an effect is distinguishable from zero. The effect size answers how big it is, and that is usually the question a reader actually has.
| Test | Effect size measure | Rough interpretation |
|---|---|---|
| t test | Standardised mean difference | 0.2 small, 0.5 medium, 0.8 large |
| Analysis of variance | Proportion of variance explained | 0.01 small, 0.06 medium, 0.14 large |
| Correlation | The coefficient itself | 0.1 small, 0.3 medium, 0.5 large |
| Chi-square | Association coefficient for tables | Interpreted relative to table size |
| Regression | Standardised coefficient; variance explained | Judged against comparable studies |
The conventional thresholds above are useful defaults and poor substitutes for judgement. Their author intended them for fields with no accumulated knowledge, and where your literature reports effect sizes for the same relationship, comparison with those is far more informative than comparison with a general convention. An effect half the size of every previous study is a finding; an effect described as medium tells the reader nothing they could not compute.
Practical significance is a separate judgement and belongs in the discussion. A correlation of 0.15 between a training intervention and productivity may be statistically significant in a large sample, may be small by any convention, and may still be worth acting on if the intervention is cheap and the outcome is valuable. Conversely a large effect on a variable nobody can influence is interesting and useless. Making this judgement explicitly, in the language of the decision it bears on, is what turns statistics into research.
Confidence intervals are more informative than p values and are increasingly expected. An interval gives the range of population values consistent with your data at a stated level of confidence, and it carries the p value's information plus the precision of the estimate. A correlation of 0.30 with an interval from 0.28 to 0.32 and one of 0.30 with an interval from 0.02 to 0.55 are both significant and are entirely different findings, and only the interval shows it.
Intervals also handle the null result well. A non-significant difference whose interval runs from a trivially small value to another trivially small value is evidence that any effect is small. A non-significant difference whose interval spans from a large negative to a large positive value is evidence of nothing at all, and calling both no effect conceals the difference between an informative null and an uninformative one.
Report, for every test: the test name, the test statistic with its degrees of freedom, the exact p value, the effect size, and where possible the confidence interval. That is one line per hypothesis and it is the reporting standard in most business journals.
القيمة الاحتمالية تجيب هل الأثر مميَّز عن الصفر. وحجم الأثر يجيب كم هو كبير، وهذا عادةً السؤال الذي يملكه القارئ فعلًا.
| الاختبار | مقياس حجم الأثر | التفسير التقريبي |
|---|---|---|
| اختبار t | فرق المتوسطات المعياري | 0.2 صغير، 0.5 متوسط، 0.8 كبير |
| تحليل التباين | نسبة التباين المفسَّر | 0.01 صغير، 0.06 متوسط، 0.14 كبير |
| الارتباط | المعامل نفسه | 0.1 صغير، 0.3 متوسط، 0.5 كبير |
| مربع كاي | معامل اقتران للجداول | يُفسَّر نسبةً إلى حجم الجدول |
| الانحدار | المعامل المعياري؛ والتباين المفسَّر | يُحكَم عليه مقابل دراسات مماثلة |
والعتبات المتعارفة أعلاه افتراضات نافعة وبدائل ضعيفة عن الحكم. فمؤلفها قصدها لحقول بلا معرفة متراكمة، وحيث تُبلغ أدبياتك بأحجام أثر للعلاقة نفسها، تكون المقارنة بها أكثر إفادةً بكثير من المقارنة بعُرف عام. فأثرٌ بنصف حجم كل دراسة سابقة نتيجة؛ وأثرٌ موصوف بـ«متوسط» لا يخبر القارئ شيئًا لم يكن ليحسبه.
والدلالة العملية حكمٌ منفصل ومكانه المناقشة. فارتباطٌ قدره 0.15 بين تدخّل تدريبي والإنتاجية قد يكون دالًا إحصائيًا في عيّنة كبيرة، وقد يكون صغيرًا بأي عُرف، وقد يظل يستحق التصرف بناءً عليه إن كان التدخل رخيصًا والمخرَج ثمينًا. وبالعكس، أثرٌ كبير على متغير لا يستطيع أحد التأثير فيه مثيرٌ وعديم الفائدة. واتخاذ هذا الحكم صراحةً، بلغة القرار الذي يمسّه، هو ما يحوّل الإحصاء إلى بحث.
وفترات الثقة أكثر إفادةً من القيم الاحتمالية ومتوقَّعة بازدياد. فالفترة تعطي مدى قيم المجتمع المتسقة مع بياناتك بمستوى ثقة مذكور، وهي تحمل معلومة القيمة الاحتمالية مع دقة التقدير. فارتباطٌ 0.30 بفترة من 0.28 إلى 0.32 وارتباطٌ 0.30 بفترة من 0.02 إلى 0.55 كلاهما دال وهما نتيجتان مختلفتان تمامًا، والفترة وحدها تُظهر ذلك.
والفترات تعالج النتيجة الصفرية جيدًا أيضًا. ففرقٌ غير دال تمتد فترته من قيمة صغيرة تافهة إلى أخرى صغيرة تافهة دليلٌ على أن أي أثر صغير. وفرقٌ غير دال تمتد فترته من سالب كبير إلى موجب كبير دليلٌ على لا شيء البتة، وتسمية كليهما «لا أثر» تخفي الفرق بين صفرية مفيدة وأخرى غير مفيدة.
وأبلغ، لكل اختبار: باسم الاختبار، وإحصاءة الاختبار بدرجات حريتها، والقيمة الاحتمالية الدقيقة، وحجم الأثر، وحيثما أمكن فترة الثقة. وهذا سطرٌ واحد لكل فرضية وهو معيار الإبلاغ في معظم مجلات الأعمال.
Most inferential errors in business theses are errors of language rather than of computation. The statistics are right and the sentence claims more than the statistics support.
Causal language from correlational designs. Increases, drives, leads to, improves, and causes all assert an ordering that a cross-sectional design cannot establish. The available verbs are is associated with, is related to, predicts in the statistical sense, and varies with. This is the single most frequent overreach in the field, and it is caught by searching your results chapter for causal verbs and checking each against the design.
Accepting the null. No significant relationship was found is correct; there is no relationship is not. The second claims to have proven absence, which no test does. The distinction matters most in the discussion, where an unsupported hypothesis is frequently written up as a positive finding of no effect.
Confusing statistical with practical significance. A significant result is not automatically an important one, and in a sample of a thousand, trivial differences reach significance routinely. State the effect size, and where it is small, say so rather than letting the word significant do work it was not intended for.
Selective reporting. Running many tests and reporting the significant ones inflates the error rate invisibly. Report every test specified in the methodology, whatever it produced. Where you ran exploratory tests not planned in advance, label them exploratory and note that their p values are not protected against multiple comparisons.
Interpreting a coefficient without its context. A regression coefficient is the estimated change in the outcome per unit change in the predictor, holding the other predictors constant. That last clause is essential and is routinely dropped, and it matters because the coefficient's meaning depends entirely on which other variables are in the model.
The structure of a results section follows from this. Restate the hypothesis. Report the descriptive picture. Report the test with statistic, degrees of freedom, exact p value, effect size, and interval. State in one sentence whether the hypothesis was supported. Then stop, because interpretation belongs in the discussion, and mixing the two makes it impossible for a reader who disagrees with your interpretation to accept your evidence.
One final discipline. Before submitting, read every sentence in the results chapter and ask what design would be required to support it. Any sentence requiring a design you did not use should be rewritten, and doing this pass will typically change eight or ten sentences in a way no reader ever notices, which is exactly the point.
معظم الأخطاء الاستدلالية في أطروحات الأعمال أخطاء لغة لا حساب. فالإحصاءات صحيحة والجملة تدّعي أكثر مما تسنده الإحصاءات.
اللغة السببية من تصاميم ارتباطية. فـ«يزيد» و«يدفع» و«يؤدي إلى» و«يحسّن» و«يسبّب» كلها تجزم بترتيب لا يستطيع تصميم مقطعي إثباته. والأفعال المتاحة «يرتبط بـ» و«يتصل بـ» و«يتنبأ» بالمعنى الإحصائي و«يتباين مع». وهذا أشيع تجاوز في الحقل، ويُلتقَط بالبحث في فصل نتائجك عن الأفعال السببية وفحص كلٍّ مقابل التصميم.
قبول الصفرية. فـ«لم يُوجَد ارتباط دال» صحيح؛ و«لا توجد علاقة» ليس كذلك. فالثانية تدّعي البرهنة على الغياب، وهذا لا يفعله أي اختبار. والتمييز يهم أكثر ما يهم في المناقشة، حيث تُكتَب فرضيةٌ غير مسنودة كثيرًا نتيجةً إيجابية بانعدام الأثر.
الخلط بين الدلالة الإحصائية والعملية. فالنتيجة الدالة ليست تلقائيًا مهمة، وفي عيّنة من ألف تبلغ الفروق التافهة الدلالة روتينيًا. اذكر حجم الأثر، وحيث يكون صغيرًا فقُل ذلك بدل ترك كلمة «دال» تؤدي عملًا لم تُقصَد له.
الإبلاغ الانتقائي. فتشغيل اختبارات كثيرة والإبلاغ بالدالة منها يضخّم معدل الخطأ بلا رؤية. أبلغ بكل اختبار محدَّد في المنهجية، أيًّا كان ما أنتجه. وحيث شغّلت اختبارات استكشافية غير مخطَّطة سلفًا، فسمِّها استكشافية ودوّن أن قيمها الاحتمالية غير محميّة من المقارنات المتعددة.
تفسير معامل بلا سياقه. فمعامل الانحدار هو التغير المقدَّر في المخرَج لكل وحدة تغير في المتنبئ، مع تثبيت المتنبئات الأخرى. وتلك العبارة الأخيرة جوهرية وتُسقَط روتينيًا، وهي تهم لأن معنى المعامل يعتمد كليًا على أي المتغيرات الأخرى في النموذج.
وبنية قسم النتائج تتبع من هذا. أعد ذكر الفرضية. وأبلغ بالصورة الوصفية. وأبلغ بالاختبار بإحصاءته ودرجات حريته وقيمته الاحتمالية الدقيقة وحجم الأثر والفترة. واذكر في جملة هل سُنِدت الفرضية. ثم توقّف، لأن التفسير مكانه المناقشة، وخلط الاثنين يجعل من المستحيل على قارئ يخالف تفسيرك أن يقبل أدلتك.
وانضباط أخير. قبل التسليم، اقرأ كل جملة في فصل النتائج واسأل أي تصميم يلزم لإسنادها. وكل جملة تتطلب تصميمًا لم تستخدمه ينبغي إعادة كتابتها، وإجراء هذا المرور سيغيّر عادةً ثماني أو عشر جمل بطريقة لا يلاحظها أي قارئ، وهذا بالضبط المقصود.