Descriptive statistics summarise what is in a dataset. They make no claim beyond the cases you observed, which is exactly why they are the necessary first stage: you cannot sensibly test a relationship between two variables you have not looked at.
The distinction that organizes the field is simple. Descriptive statistics describe the sample. The mean tenure of these two hundred respondents is 4.3 years, and that statement is exactly true of these two hundred people and makes no claim about anyone else. Inferential statistics generalise to the population and carry uncertainty as a result. Descriptive statistics are true by construction; inferential ones are probabilistic, which is why they are reported with intervals and p values and descriptive ones are not.
Most students under-report descriptives, treating them as preliminary material to be got through before the real analysis. That is a mistake for three reasons. Descriptives are how you find the errors that survived data preparation. They are how a reader judges whether your sample resembles the population you claim it represents. And a substantial number of business research questions are genuinely descriptive, so for those studies this is not preliminary work but the finding itself.
This framework covers the four questions descriptives answer, the three measures of central tendency and which the data permit, dispersion and why it matters as much as the mean, distribution shape, bivariate description, and how to present all of it. It is my own synthesis, written in my own words and grounded in recognized scholarship.
الإحصاء الوصفي يلخّص ما في مجموعة بيانات. ولا يطلق ادّعاءً يتجاوز الحالات التي لاحظتها، وهذا بالضبط لماذا هو المرحلة الأولى اللازمة: فلا يمكنك اختبار علاقة بين متغيرين لم تنظر فيهما بشكل معقول.
والتمييز الذي ينظّم الحقل بسيط. الإحصاء الوصفي يصف العيّنة. فمتوسط مدة خدمة هؤلاء المئتي مستجيب 4.3 سنوات، وتلك العبارة صحيحة بالضبط عن هؤلاء المئتين ولا تطلق ادّعاءً عن أي أحد آخر. والإحصاء الاستدلالي يعمّم على المجتمع ويحمل عدم يقين نتيجةً لذلك. فالإحصاء الوصفي صحيح بالبناء؛ والاستدلالي احتمالي، ولهذا يُبلَّغ عنه بفترات وقيم احتمالية ولا يُبلَّغ عن الوصفي كذلك.
ومعظم الطلاب يقصّرون في الإبلاغ عن الوصفيات، معاملين إياها مادةً تمهيدية تُجتاز قبل التحليل الحقيقي. وهذا خطأ لثلاثة أسباب. فالوصفيات هي كيف تجد الأخطاء التي نجت من إعداد البيانات. وهي كيف يحكم القارئ هل تشبه عيّنتك المجتمع الذي تدّعي أنها تمثّله. وعددٌ كبير من أسئلة بحوث الأعمال وصفي فعلًا، فلتلك الدراسات هذا ليس عملًا تمهيديًا بل النتيجة نفسها.
ويغطي هذا الإطار الأسئلة الأربعة التي تجيبها الوصفيات، ومقاييس النزعة المركزية الثلاثة وأيها تجيزه البيانات، والتشتت ولماذا يهم بقدر المتوسط، وشكل التوزيع، والوصف الثنائي، وكيف يُعرَض هذا كله. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
A single variable is described by answering four questions. Reporting the mean alone answers one of them, which is why a table of means with no dispersion tells a reader very little.
Frequency is the count and percentage in each category. It is the whole of the description for a nominal variable, and it is what demographic tables report. Give both the count and the percentage, since a percentage without its base is uninterpretable and a count without a percentage is hard to compare across groups.
Central tendency locates the middle. Which measure is legitimate depends on the measurement level, and this is the most frequently violated rule in student analysis: a mean computed on a nominal variable is arithmetic performed on category codes and means nothing whatever.
Dispersion describes how spread out the values are, and it carries at least as much information as the middle. Two departments with the same mean satisfaction score are very different places if one has a standard deviation of 0.3 and the other of 1.4: the first is uniformly moderate and the second contains both very satisfied and very dissatisfied people. A finding of no difference in means between groups is frequently accompanied by a large difference in variance, which is a finding in its own right and is invisible if only means are reported.
Shape concerns whether the distribution is symmetric or skewed and whether it is flatter or more peaked than a normal curve. Skewness matters practically because it determines whether the mean is a fair summary: in a right-skewed distribution such as income or firm size, the mean sits well above the typical value and reporting it alone misleads. This is why income and revenue are conventionally described with the median.
The practical rule is to report all four for every variable that matters to your argument. For continuous variables that is the mean, the standard deviation, the minimum, and the maximum at minimum, with the median added wherever the distribution is skewed. For categorical variables it is the count and percentage per category. Anything less forces the reader to trust a summary they cannot inspect.
المتغير المفرد يُوصَف بإجابة أربعة أسئلة. والإبلاغ بالمتوسط وحده يجيب واحدًا منها، ولهذا فجدولُ متوسطات بلا تشتت يخبر القارئ قليلًا جدًا.
والتكرار هو العدد والنسبة في كل فئة. وهو الوصف كله لمتغير اسمي، وهو ما تُبلغ به الجداول الديموغرافية. أعطِ العدد والنسبة معًا، لأن النسبة بلا أساسها غير قابلة للتفسير والعدد بلا نسبة صعب المقارنة عبر المجموعات.
والنزعة المركزية تحدد موقع الوسط. وأي مقياس مشروع يعتمد على مستوى القياس، وهذه أكثر القواعد انتهاكًا في تحليل الطلاب: فمتوسطٌ محسوب على متغير اسمي حسابٌ يُجرى على رموز فئات ولا يعني شيئًا البتة.
والتشتت يصف كم تنتشر القيم، ويحمل معلومة لا تقل عن الوسط. فقسمان بمتوسط درجة رضا واحد مكانان مختلفان جدًا إن كان لأحدهما انحراف معياري 0.3 وللآخر 1.4: فالأول معتدل بانتظام والثاني يحوي راضين جدًا وغير راضين جدًا معًا. ونتيجةُ «لا فرق في المتوسطات بين المجموعات» كثيرًا ما يرافقها فرق كبير في التباين، وهذه نتيجة بذاتها وهي غير مرئية إن لم يُبلَّغ إلا بالمتوسطات.
والشكل يخص هل التوزيع متماثل أم ملتوٍ وهل هو أبطح أم أشد تدبّبًا من المنحنى الاعتدالي. والالتواء يهم عمليًا لأنه يحدد هل المتوسط ملخصٌ منصف: ففي توزيع ملتوٍ يمينًا كالدخل أو حجم المنشأة، يجلس المتوسط فوق القيمة النموذجية بكثير والإبلاغ به وحده مضلِّل. ولهذا يُوصَف الدخل والإيراد عرفًا بالوسيط.
والقاعدة العملية أن تُبلغ بالأربعة كلها لكل متغير يهم حجتك. فللمتغيرات المتصلة هذا المتوسط والانحراف المعياري والأدنى والأقصى كحد أدنى، مع إضافة الوسيط حيثما كان التوزيع ملتويًا. وللمتغيرات الفئوية هو العدد والنسبة لكل فئة. وأي أقل من ذلك يُجبر القارئ على الوثوق بملخص لا يستطيع فحصه.
Three measures locate the middle and three describe the spread. Choosing among them is decided by the measurement level and by the shape of the distribution.
The mean uses every value and is the basis of most inferential statistics, which is why it is the default for interval and ratio data. Its weakness is sensitivity to extremes: one firm with revenue a hundred times the others moves the mean substantially and the median not at all. The median is the middle value when the data are ordered, is legitimate for ordinal data, and is the honest summary for any skewed distribution. The mode is the most frequent value and is the only central tendency available for nominal data.
The practical rule for business data: report the mean for scale variables that are roughly symmetric, and report both the mean and the median for anything involving money, size, tenure, or counts, because those variables are almost always right-skewed. Where the two differ substantially, say so, since the difference itself tells the reader about the distribution.
For dispersion, three measures matter. The range, meaning maximum minus minimum, is the crudest and is most useful as a check on data errors rather than as a summary. The interquartile range, the distance between the twenty-fifth and seventy-fifth percentiles, describes the middle half of the data and is the companion to the median. The standard deviation is the average distance of values from the mean and is the companion to the mean, and is also what most inferential procedures are built on.
Interpreting a standard deviation requires context, because its size depends on the scale's units. On a five-point agreement scale, a standard deviation around 0.5 indicates high agreement among respondents, around 1.0 is typical, and above 1.3 indicates a genuinely divided sample, which is often more interesting than the mean. For unbounded variables such as revenue, the coefficient of variation, meaning the standard deviation divided by the mean, allows comparison of variability across variables measured in different units.
Two habits improve every descriptive section. Report central tendency and dispersion together, always, since either alone is half a description. And where you compare groups, report both for each group rather than only the means, because two groups can share a mean and differ entirely in how their members are distributed around it.
ثلاثة مقاييس تحدد الوسط وثلاثة تصف الانتشار. والاختيار بينها يقرره مستوى القياس وشكل التوزيع.
والمتوسط يستخدم كل قيمة وهو أساس معظم الإحصاء الاستدلالي، ولهذا هو الافتراضي للبيانات الفترية والنسبية. وضعفه الحساسية للمتطرفات: فمنشأة واحدة بإيراد مئة ضعف الأخريات تحرّك المتوسط كثيرًا ولا تحرّك الوسيط البتة. والوسيط القيمةُ الوسطى حين تُرتَّب البيانات، وهو مشروع للبيانات الرتبية، وهو الملخص الصادق لأي توزيع ملتوٍ. والمنوال القيمةُ الأكثر تكرارًا وهو النزعة المركزية الوحيدة المتاحة للبيانات الاسمية.
والقاعدة العملية لبيانات الأعمال: أبلغ بالمتوسط للمتغيرات المقياسية المتماثلة تقريبًا، وأبلغ بالمتوسط والوسيط معًا لأي شيء يتعلق بالمال أو الحجم أو مدة الخدمة أو الأعداد، لأن تلك المتغيرات ملتوية يمينًا شبه دائمًا. وحيث يختلف الاثنان كثيرًا، فقُل ذلك، لأن الفرق نفسه يخبر القارئ عن التوزيع.
وللـتشتت، ثلاثة مقاييس تهم. فـالمدى، أي الأقصى ناقص الأدنى، أخشنها وأنفعها فحصًا لأخطاء البيانات لا ملخصًا. والمدى الربيعي، المسافة بين المئين الخامس والعشرين والخامس والسبعين، يصف النصف الأوسط من البيانات وهو رفيق الوسيط. والانحراف المعياري متوسطُ مسافة القيم عن المتوسط وهو رفيق المتوسط، وهو أيضًا ما تُبنى عليه معظم الإجراءات الاستدلالية.
وتفسير الانحراف المعياري يتطلب سياقًا، لأن حجمه يعتمد على وحدات المقياس. فعلى مقياس موافقة خماسي، يشير انحراف معياري نحو 0.5 إلى اتفاق عالٍ بين المستجيبين، ونحو 1.0 نموذجي، وفوق 1.3 يشير إلى عيّنة منقسمة فعلًا، وهذا كثيرًا ما يكون أطرف من المتوسط. وللمتغيرات غير المحدودة كالإيراد، يتيح معامل الاختلاف، أي الانحراف المعياري مقسومًا على المتوسط، مقارنة التقلب عبر متغيرات مقيسة بوحدات مختلفة.
وعادتان تحسّنان كل قسم وصفي. أبلغ بالنزعة المركزية والتشتت معًا، دائمًا، لأن أيًّا منهما وحده نصفُ وصف. وحيث تقارن مجموعات، أبلغ بالاثنين لكل مجموعة لا بالمتوسطات وحدها، لأن مجموعتين قد تتشاركان متوسطًا وتختلفان كليًا في كيفية توزع أفرادهما حوله.
Shape determines which summaries are honest, and it also determines which inferential tests are available, which is why it is checked before rather than after.
| Feature | What it means | What to do |
|---|---|---|
| Right skew | A long tail of high values; mean above median | Report the median; consider a log transformation |
| Left skew | A long tail of low values; mean below median | Report the median; check for a ceiling effect |
| Ceiling or floor effect | Many cases at the top or bottom of the scale | The measure may not discriminate at that end |
| Bimodality | Two peaks in the distribution | Look for a grouping variable that explains it |
| Heavy tails | More extreme values than a normal curve | Robust methods, or report with and without extremes |
Bimodality is the most informative shape and the most often missed, because a mean computed on a bimodal distribution describes a value that almost nobody holds. Two peaks usually mean two populations are mixed, and finding the variable that separates them, often a department, a role, or a tenure band, converts an odd-looking histogram into a substantive finding.
Ceiling effects matter particularly in business surveys of attitudes, where items about a socially approved practice frequently produce distributions crushed against the top of the scale. When eighty percent of respondents choose four or five, the measure cannot distinguish among them, and a null result for that variable may reflect the measure rather than the world.
Describing pairs extends the same logic to two variables. For two categorical variables, a cross-tabulation with row or column percentages is the description; choose the percentage direction that matches the question, since the percentage of women who are managers and the percentage of managers who are women are different numbers answering different questions.
For two continuous variables, a scatterplot is the description and should be produced before any correlation is computed. The correlation coefficient summarises a linear relationship in a single number, and it reports a curved relationship as weak, an outlier-driven relationship as strong, and two distinct subgroups as a moderate association that describes neither. All three are visible in a scatterplot and invisible in the coefficient.
For one categorical and one continuous variable, the description is the continuous variable's central tendency and dispersion within each category, presented as a table or a boxplot. This is the descriptive form of a group comparison, and it should always precede the significance test, because a difference of 0.1 on a five-point scale can be statistically significant in a large sample and is not worth discussing.
الشكل يحدد أي الملخصات صادقة، ويحدد أيضًا أي الاختبارات الاستدلالية متاحة، ولهذا يُفحَص قبل لا بعد.
| السمة | ما تعنيه | ما يُفعَل |
|---|---|---|
| التواء يميني | ذيل طويل من القيم العالية؛ المتوسط فوق الوسيط | أبلغ بالوسيط؛ وفكّر في تحويل لوغاريتمي |
| التواء يساري | ذيل طويل من القيم المنخفضة؛ المتوسط دون الوسيط | أبلغ بالوسيط؛ وافحص أثر السقف |
| أثر السقف أو الأرضية | حالات كثيرة عند أعلى المقياس أو أدناه | قد لا يميّز المقياس عند ذلك الطرف |
| ثنائية المنوال | قمتان في التوزيع | ابحث عن متغير تجميع يفسّرها |
| ذيول ثقيلة | قيم متطرفة أكثر من منحنى اعتدالي | طرق متينة، أو الإبلاغ مع المتطرفات وبدونها |
وثنائية المنوال أكثر الأشكال إفادةً وأكثرها إغفالًا، لأن متوسطًا محسوبًا على توزيع ثنائي المنوال يصف قيمةً لا يحملها أحد تقريبًا. والقمتان تعنيان عادةً أن مجتمعين مختلطان، وإيجادُ المتغير الذي يفصلهما، وهو غالبًا قسم أو دور أو فئة مدة خدمة، يحوّل رسمًا بيانيًا غريب المظهر إلى نتيجة جوهرية.
وآثار السقف تهم خاصةً في مسوح الأعمال للاتجاهات، حيث تُنتج البنود عن ممارسة مستحسَنة اجتماعيًا توزيعات مضغوطة على أعلى المقياس. فحين يختار ثمانون بالمئة من المستجيبين أربعة أو خمسة، لا يستطيع المقياس التمييز بينهم، وقد تعكس نتيجةٌ صفرية لذلك المتغير المقياسَ لا العالم.
ووصف الأزواج يمدّ المنطق نفسه إلى متغيرين. فلمتغيرين فئويين، يكون الجدول المتقاطع بنسب الصفوف أو الأعمدة هو الوصف؛ واختر اتجاه النسبة الذي يطابق السؤال، لأن نسبة النساء اللواتي هن مديرات ونسبة المديرين الذين هم نساء رقمان مختلفان يجيبان سؤالين مختلفين.
ولمتغيرين متصلين، يكون رسم الانتشار هو الوصف وينبغي إنتاجه قبل حساب أي ارتباط. فمعامل الارتباط يلخّص علاقة خطية في رقم واحد، وهو يُبلغ عن علاقة منحنية ضعيفةً، وعن علاقة تقودها قيمة شاذة قويةً، وعن فئتين فرعيتين متمايزتين ارتباطًا متوسطًا لا يصف أيًّا منهما. والثلاثة كلها مرئية في رسم الانتشار وغير مرئية في المعامل.
ولمتغير فئوي وآخر متصل، يكون الوصف نزعةَ المتغير المتصل المركزية وتشتته داخل كل فئة، معروضًا جدولًا أو رسمًا صندوقيًا. وهذه الصيغة الوصفية لمقارنة مجموعات، وينبغي أن تسبق اختبار الدلالة دائمًا، لأن فرقًا قدره 0.1 على مقياس خماسي قد يكون دالًا إحصائيًا في عيّنة كبيرة ولا يستحق النقاش.
The descriptive section of a results chapter has a conventional structure. Following it lets a reader find what they are looking for without hunting.
Start with the response and sample profile. How many were approached, how many responded, the response rate, and how many cases entered analysis after screening. Then the demographic composition as a table of counts and percentages, and, where possible, a comparison with the population, since a sample that matches the population on observable characteristics is evidence against non-response bias and takes one row of a table to demonstrate.
Then the main variables. One table with a row per variable and columns for the number of valid cases, the mean, the standard deviation, the minimum, and the maximum, with the median added for skewed variables and the scale reliability where relevant. This single table is often the most-consulted page in a thesis, since every later result is interpreted against it.
Then the correlation matrix among the main variables, which serves two purposes: it is the bivariate description, and it is the diagnostic that shows whether any predictors are collinear. Conventional practice is to report the correlations below the diagonal, with significance marked, and to place the mean and standard deviation in the first two columns so the table is self-contained.
Then group comparisons where your questions involve them, descriptively before inferentially: means and standard deviations by group, so the reader sees the size of the difference before being told whether it is statistically significant.
Four reporting conventions are worth following. Use a consistent number of decimal places, usually two for means and standard deviations and two or three for correlations, since varying precision across a table looks careless. Report exact percentages with their base. Do not report a mean to three decimals from a five-point scale, which implies a precision the measurement does not have. And describe in the text what the table shows, in two or three sentences pointing at what matters, because a table presented without interpretation has handed the reader your job.
Finally, keep the description free of inference. The descriptive section reports what is in the sample; it does not say that a difference is significant, that a relationship holds in the population, or that one group is better than another. Those are claims for the sections that follow, and mixing them into the description is the most common structural error in a results chapter.
القسم الوصفي من فصل النتائج له بنية متعارفة. واتباعها يتيح للقارئ إيجاد ما يبحث عنه دون تنقيب.
ابدأ بالاستجابة وملف العيّنة. كم جرى التواصل معهم، وكم استجابوا، ومعدل الاستجابة، وكم حالةً دخلت التحليل بعد الفرز. ثم التركيب الديموغرافي جدولَ أعداد ونسب، وحيثما أمكن، مقارنةً بالمجتمع، لأن عيّنةً تطابق المجتمع في الخصائص الملحوظة دليلٌ ضد تحيّز عدم الاستجابة ويستغرق إظهاره صفًا واحدًا من جدول.
ثم المتغيرات الرئيسة. جدولٌ واحد بصف لكل متغير وأعمدة لعدد الحالات الصالحة والمتوسط والانحراف المعياري والأدنى والأقصى، مع إضافة الوسيط للمتغيرات الملتوية والثبات حيثما كان ذا صلة. وهذا الجدول الواحد كثيرًا ما يكون أكثر صفحات الأطروحة مراجعةً، لأن كل نتيجة لاحقة تُفسَّر مقابله.
ثم مصفوفة الارتباط بين المتغيرات الرئيسة، وهي تخدم غرضين: فهي الوصف الثنائي، وهي التشخيص الذي يُظهر هل أي متنبئات متعددة الخطية. والممارسة المتعارفة الإبلاغُ بالارتباطات تحت القطر، مع تعليم الدلالة، ووضعُ المتوسط والانحراف المعياري في العمودين الأولين ليكون الجدول قائمًا بذاته.
ثم مقارنات المجموعات حيث تتضمنها أسئلتك، وصفيًا قبل استدلاليًا: متوسطات وانحرافات معيارية بحسب المجموعة، ليرى القارئ حجم الفرق قبل أن يُخبَر هل هو دال إحصائيًا.
وأربعة أعراف إبلاغ تستحق الاتباع. استخدم عددًا متسقًا من المنازل العشرية، منزلتين للمتوسطات والانحرافات ومنزلتين أو ثلاثًا للارتباطات عادةً، لأن تباين الدقة عبر جدول يبدو مهملًا. وأبلغ بـنسب دقيقة بأساسها. ولا تُبلغ بمتوسط بثلاث منازل عشرية من مقياس خماسي، فهذا يوحي بدقة لا يملكها القياس. وصِف في النص ما يُظهره الجدول، في جملتين أو ثلاث تشير إلى ما يهم، لأن جدولًا معروضًا بلا تفسير قد سلّم القارئ وظيفتك.
وأخيرًا، أبقِ الوصف خاليًا من الاستدلال. فالقسم الوصفي يُبلغ بما في العيّنة؛ ولا يقول إن فرقًا دال، ولا إن علاقةً تصح في المجتمع، ولا إن مجموعةً خير من أخرى. فتلك ادّعاءات للأقسام التالية، وخلطها في الوصف أشيع خطأ بنيوي في فصل نتائج.