Saif Ali AlghamdiTransformation & Growth Advisor
تواصل
Foundation of Business Researchأسس بحوث الأعمالResults and discussionالنتائج والمناقشة
RESULTS AND DISCUSSION · PhDالنتائج والمناقشة · دكتوراه

Preparing the Dataإعداد البيانات

SectionالقسمResults and discussionالنتائج والمناقشة
Reading timeزمن القراءة11 min١١ دقيقة
ByإعدادSaif Alghamdiسيف الغامدي
One

Overview

Topic: Preparing raw data for analysis
Covers: Screening cases, missing values, recoding, assumption checks, and documentation
Use: Producing an analysis file you can defend and reproduce
Level: Doctoral business research
By: Saif Alghamdi

Between collecting data and analysing it sits a stage that gets almost no attention and determines much of what follows. Raw data are never analysable as they arrive: some responses are not real, some values are missing, some items need reversing, and some variables do not exist yet.

The stage matters for two reasons. The first is correctness: an analysis run on badly prepared data produces results that are precisely computed and wrong, and no reviewer can see the error because the error is upstream of everything reported. The second is defensibility: every preparation decision changes the dataset, and a decision made without a rule looks, in retrospect, like a decision made because of what it did to the results.

That second point is the governing principle of this whole stage. Decide the rules before looking at the outcomes. Whether to drop cases below a completion threshold, how to treat outliers, and how to handle missing values are all decisions that can move a coefficient across a significance threshold, and a researcher who chooses among them after seeing that effect is no longer testing a hypothesis. Writing the rules into the methodology in advance removes the problem entirely.

This framework covers the six steps in order, case screening and what counts as a non-response, the three missing data mechanisms and what each permits, recoding and scale construction, the assumption checks that precede any test, and how to document all of it. It is my own synthesis, written in my own words and grounded in recognized scholarship.

Note: Never edit the raw file. Keep it read-only, do all preparation in a script or a documented sequence of steps, and generate the analysis file from the raw one. If you cannot rebuild your analysis file from the raw data, you cannot answer the question of what you did to it.
الأول

نظرة عامة

الموضوع: إعداد البيانات الخام للتحليل
يغطّي: فرز الحالات، والقيم المفقودة، وإعادة الترميز، وفحوص الافتراضات، والتوثيق
الاستخدام: إنتاج ملف تحليل تستطيع الدفاع عنه وإعادة إنتاجه
المستوى: بحوث الأعمال لمرحلة الدكتوراه
إعداد: سيف الغامدي

بين جمع البيانات وتحليلها تجلس مرحلة لا تكاد تنال انتباهًا وتحدد كثيرًا مما يليها. فالبيانات الخام ليست قابلة للتحليل كما تصل أبدًا: فبعض الاستجابات ليست حقيقية، وبعض القيم مفقودة، وبعض البنود يحتاج عكسًا، وبعض المتغيرات لم يوجد بعد.

وتهم المرحلة لسببين. الأول الصحة: فتحليل يُشغَّل على بيانات سيئة الإعداد يُنتج نتائج محسوبة بدقة وخاطئة، ولا يستطيع محكّم رؤية الخطأ لأن الخطأ أعلى مجرى كل ما يُبلَّغ عنه. والثاني القابلية للدفاع: فكل قرار إعداد يغيّر مجموعة البيانات، والقرار المتخذ بلا قاعدة يبدو، بأثر رجعي، قرارًا اتُّخذ بسبب ما فعله بالنتائج.

وتلك النقطة الثانية هي المبدأ الحاكم لهذه المرحلة كلها. قرّر القواعد قبل النظر في المخرجات. فهل تُسقط الحالات دون عتبة إكمال، وكيف تُعامَل القيم الشاذة، وكيف يُعالَج المفقود، كلها قرارات تستطيع تحريك معامل عبر عتبة دلالة، والباحث الذي يختار بينها بعد رؤية ذلك الأثر لم يعد يختبر فرضية. وكتابةُ القواعد في المنهجية سلفًا تزيل المشكلة تمامًا.

ويغطي هذا الإطار الخطوات الست بالترتيب، وفرز الحالات وما يُعدّ عدم استجابة، وآليات الفقدان الثلاث وما يجيزه كلٌّ، وإعادة الترميز وبناء المقاييس، وفحوص الافتراضات السابقة لأي اختبار، وكيف يُوثَّق هذا كله. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.

ملاحظة: لا تحرر الملف الخام أبدًا. أبقِه للقراءة فقط، وأجرِ كل الإعداد في نص برمجي أو تسلسل خطوات موثَّق، وولّد ملف التحليل من الخام. فإن لم تستطع إعادة بناء ملف تحليلك من البيانات الخام، فلا تستطيع إجابة سؤال ماذا فعلت به.
Two

The Six Steps in Order

The order is not arbitrary. Each step assumes the previous one is done, and running them out of order produces errors that are hard to trace afterwards.

1Import and inspectget the raw file in, and look at it before touching it2Screen casesremove responses that are not real data3Handle missing valuesdecide a rule, apply it, and report the effect4Recode and scorereverse items, compute scale means, build derived variables5Check assumptionsdistributions, outliers, and what the planned tests require6Freeze and documentlock the analysis file and record every step takensix steps, in this order; the last one is what makes the first five defensible

Import and inspect means getting the file into your software and looking at it before doing anything. Check that the number of rows matches what the collection platform reported, that every variable imported with the right type, that string fields did not become numbers or the reverse, and that the value labels survived the import. Print the minimum, maximum, and number of distinct values for every variable, and look for impossible values: a tenure of ninety-four years, a scale response of 7 on a five-point item, an age of 3.

Screen cases removes responses that are not data. Recode and score constructs the variables you will actually analyse. Check assumptions establishes whether your planned tests are appropriate, and if they are not, that is a finding about the analysis plan rather than a licence to run them anyway.

Freeze and document is the step that makes the rest defensible. When preparation is finished, save the analysis file with a version and a date, and stop changing it. Every subsequent result is then traceable to one file. If a later problem requires a change, make it in the preparation script, regenerate the file with a new version number, and rerun the analysis rather than editing the file in place.

Two practices support all six. Work in a script rather than by clicking, in whatever software you use, because a script is a record of what you did that a sequence of menu selections is not. And keep a preparation log alongside it, a plain document recording each decision, its rule, and its effect on the case count, because the script records what you did and the log records why.

Note: The case count after every step is the arithmetic a reader will check. Start with the number of raw responses and end with the analysis sample, accounting for every case removed along the way, exactly as a flow diagram does for a literature search.
الثاني

الخطوات الست بالترتيب

الترتيب ليس اعتباطيًا. فكل خطوة تفترض إنجاز سابقتها، وتشغيلها خارج الترتيب يُنتج أخطاءً يصعب تتبّعها لاحقًا.

١الاستيراد والفحصأدخل الملف الخام، وانظر فيه قبل لمسه٢فرز الحالاتأزل الاستجابات التي ليست بيانات حقيقية٣معالجة المفقودقرّر قاعدة، وطبّقها، وأبلغ بالأثر٤إعادة الترميز والتسجيلاعكس البنود، واحسب متوسطات المقاييس، وابنِ المتغيرات المشتقة٥فحص الافتراضاتالتوزيعات والقيم الشاذة وما تتطلبه الاختبارات المزمعة٦التجميد والتوثيقأقفل ملف التحليل وسجّل كل خطوة اتُّخذتست خطوات بهذا الترتيب؛ والأخيرة هي ما يجعل الخمس الأولى قابلة للدفاع

والاستيراد والفحص يعني إدخال الملف إلى برمجيتك والنظر فيه قبل فعل أي شيء. افحص أن عدد الصفوف يطابق ما أبلغت به منصة الجمع، وأن كل متغير استُورد بالنوع الصحيح، وأن الحقول النصية لم تصر أرقامًا ولا العكس، وأن تسميات القيم نجت من الاستيراد. واطبع الأدنى والأقصى وعدد القيم المتمايزة لكل متغير، وابحث عن قيم مستحيلة: مدة خدمة أربع وتسعين سنة، واستجابة 7 على بند خماسي، وعمر 3.

وفرز الحالات يزيل الاستجابات التي ليست بيانات. وإعادة الترميز والتسجيل تبني المتغيرات التي ستحللها فعلًا. وفحص الافتراضات يثبّت هل الاختبارات المزمعة ملائمة، وإن لم تكن، فتلك نتيجة عن خطة التحليل لا إجازة لتشغيلها على أي حال.

والتجميد والتوثيق هي الخطوة التي تجعل الباقي قابلًا للدفاع. فحين ينتهي الإعداد، احفظ ملف التحليل بإصدار وتاريخ، وتوقف عن تغييره. وكل نتيجة لاحقة تصير عندئذ قابلة للتتبع إلى ملف واحد. وإن تطلبت مشكلة لاحقة تغييرًا، فأجرِه في نص الإعداد، وأعد توليد الملف برقم إصدار جديد، وأعد تشغيل التحليل بدل تحرير الملف في موضعه.

وممارستان تسندان الست جميعًا. اعمل في نص برمجي لا بالنقر، بأي برمجية تستخدم، لأن النص سجلٌّ لما فعلت وتسلسل اختيارات القوائم ليس كذلك. واحتفظ بـسجل إعداد إلى جانبه، مستندٍ بسيط يسجّل كل قرار وقاعدته وأثره في عدد الحالات، لأن النص يسجّل ماذا فعلت والسجل يسجّل لماذا.

ملاحظة: عدد الحالات بعد كل خطوة هو الحساب الذي سيفحصه القارئ. ابدأ بعدد الاستجابات الخام وانتهِ بعيّنة التحليل، محاسبًا عن كل حالة أُزيلت في الطريق، تمامًا كما يفعل مخطط التدفق لبحث في الأدبيات.
Three

Screening Cases and Outliers

Some responses in every dataset are not data. Removing them improves the analysis; removing the wrong ones biases it. The difference is having a rule stated in advance.

ProblemHow to detect itStandard rule
Incomplete responsesProportion of items answeredDrop below a stated completion threshold
Straight-liningZero or near-zero variance across a long blockDrop where variance is zero across many items
Impossibly fast completionDuration well below the pilot medianDrop below a stated fraction of median time
Failed attention checkAn item instructing a specific responseDrop on failure, if planned in advance
Duplicate submissionsIdentical response patterns and timestampsKeep the first, drop the rest
Ineligible respondentsA screening question at the startDrop, and report separately from non-response

State the thresholds in the methodology before collection. Below eighty percent completion and below one third of the median completion time are common and defensible choices, but any threshold is defensible if it was fixed in advance and applied consistently. Report how many cases each rule removed, and check whether removed cases differ systematically from retained ones, because a rule that disproportionately removes one department or one seniority level has changed your sample.

Outliers are a separate matter and are frequently mishandled. An outlier is an extreme value, and extreme is not the same as wrong. Three kinds occur and each has a different response. An error outlier is a data entry or coding mistake, and it should be corrected against the source or removed. An ineligible outlier comes from a case that does not belong in the population, such as a firm ten times larger than the size band you sampled, and it should be removed with the reason stated. A legitimate extreme value is a real observation from a real member of your population, and removing it because it is inconvenient is data manipulation.

For legitimate extremes, the defensible options are to retain them and report results with and without them, to use an analysis technique robust to extreme values, or to transform the variable. Reporting both sets of results is the strongest option because it lets a reader see how much the finding depends on a small number of cases, and if the finding disappears without them, that is itself important information.

Detect outliers with the methods appropriate to the analysis. For a single variable, values beyond about three standard deviations from the mean, or beyond one and a half times the interquartile range from the quartiles, are the conventional flags. For a regression, influence measures identify cases that change the coefficients disproportionately, and those are the ones that matter, since an extreme value on one variable that does not affect the model is not a problem.

Note: Never remove a case because it weakens a result. If you find yourself considering it, that is the moment the rule you should have written in advance would have protected you, and the correct action now is to report the result with the case included.
الثالث

فرز الحالات والقيم الشاذة

بعض الاستجابات في كل مجموعة بيانات ليست بيانات. وإزالتها تحسّن التحليل؛ وإزالة الخاطئة تحيّزه. والفرق وجودُ قاعدة مذكورة سلفًا.

المشكلةكيف تُكتشَفالقاعدة المعيارية
استجابات ناقصةنسبة البنود المُجابةأسقط ما دون عتبة إكمال مذكورة
الاستجابة المستقيمةتباين صفري أو شبه صفري عبر كتلة طويلةأسقط حيث التباين صفر عبر بنود كثيرة
إكمال سريع مستحيلمدة أدنى بكثير من وسيط التجربةأسقط دون كسر مذكور من الزمن الوسيط
إخفاق فحص الانتباهبند يوجّه إلى استجابة بعينهاأسقط عند الإخفاق، إن خُطِّط سلفًا
استجابات مكررةأنماط استجابة وأختام زمنية متطابقةأبقِ الأولى وأسقط الباقي
مستجيبون غير مؤهلينسؤال فرز في البدايةأسقط، وأبلغ منفصلًا عن عدم الاستجابة

واذكر العتبات في المنهجية قبل الجمع. فدون ثمانين بالمئة إكمالًا ودون ثلث الزمن الوسيط للإكمال خياران شائعان وقابلان للدفاع، لكن أي عتبة قابلة للدفاع إن ثُبِّتت سلفًا وطُبِّقت باتساق. وأبلغ بكم حالةً أزالت كل قاعدة، وافحص هل تختلف الحالات المُزالة منهجيًا عن المبقاة، لأن قاعدةً تزيل بقدر غير متناسب قسمًا واحدًا أو مستوى درجة واحدًا قد غيّرت عيّنتك.

والقيم الشاذة مسألة منفصلة ويُساء التعامل معها كثيرًا. فالقيمة الشاذة قيمةٌ متطرفة، والمتطرف ليس الخطأ. وثلاثة أنواع تقع ولكلٍّ استجابة مختلفة. فـالشاذة الخطأ غلطُ إدخال أو ترميز، وينبغي تصحيحها مقابل المصدر أو إزالتها. والشاذة غير المؤهلة تأتي من حالة لا تنتمي إلى المجتمع، كمنشأة أكبر عشر مرات من فئة الحجم التي أخذت منها عيّنة، وينبغي إزالتها بذكر السبب. والقيمة المتطرفة المشروعة ملاحظةٌ حقيقية من عضو حقيقي في مجتمعك، وإزالتها لأنها غير مريحة تلاعبٌ بالبيانات.

وللمتطرفات المشروعة، الخيارات القابلة للدفاع إبقاؤها والإبلاغ بالنتائج معها وبدونها، أو استخدام تقنية تحليل متينة أمام القيم المتطرفة، أو تحويل المتغير. والإبلاغ بمجموعتي النتائج أقوى الخيارات لأنه يتيح للقارئ رؤية كم تعتمد النتيجة على عدد صغير من الحالات، وإن اختفت النتيجة بدونها فتلك نفسها معلومة مهمة.

واكتشف الشواذ بالطرق الملائمة للتحليل. فلمتغير مفرد، تكون القيم بعد نحو ثلاثة انحرافات معيارية عن المتوسط، أو بعد مرة ونصف المدى الربيعي عن الربيعيات، هي الأعلام المتعارَفة. وللانحدار، تحدد مقاييس التأثير الحالات التي تغيّر المعاملات بقدر غير متناسب، وتلك هي التي تهم، لأن قيمةً متطرفة على متغير لا تؤثر في النموذج ليست مشكلة.

ملاحظة: لا تزل حالةً أبدًا لأنها تُضعف نتيجة. فإن وجدت نفسك تفكر في ذلك، فتلك اللحظة التي كانت القاعدة التي كان ينبغي كتابتها سلفًا ستحميك فيها، والفعل الصحيح الآن الإبلاغُ بالنتيجة مع إدراج الحالة.
Four

Missing Values

Missing data are not a nuisance to be cleared away. The pattern of what is missing is information, and the right treatment depends entirely on why the values are absent.

Completely at randomAt randomNot at randommissingness unrelated toanythingmissingness explained by othervariables you havemissingness depends on themissing value itselfdropping cases is unbiasedimputation can work if thepredictors are in the datano technique fixes it; it must bedisclosedthree mechanisms; only the third is fatal, and only reasoning identifies it

Start by quantifying: the percentage missing per variable and per case, and whether missingness clusters. A variable missing on twenty percent of cases is a problem; a variable missing on two percent is not. Cases missing many values are usually better dropped than repaired.

Then reason about the mechanism, which cannot be determined statistically and must be argued. If income is missing more often among higher earners, missingness depends on the missing value itself and is not random in the technical sense. If a block of items is missing for everyone who took the survey on a mobile device, missingness is explained by a variable you have, and can be handled. If values are missing because of a transmission error, missingness is unrelated to anything.

The treatments follow from the mechanism. Listwise deletion, meaning dropping any case with a missing value on the analysis variables, is the default in most software and is unbiased only when data are missing completely at random. Its real cost is sample size: with ten variables each missing five percent, listwise deletion can remove a third of the sample. Pairwise deletion uses all available data for each computation and produces a correlation matrix in which different cells rest on different cases, which causes problems in multivariate techniques.

Mean substitution, replacing a missing value with the variable's mean, is common in student work and is not recommended: it reduces variance artificially and weakens the very relationships you are testing. Regression or multiple imputation estimates missing values from the other variables and is the preferred modern approach when missingness is explained by variables you have. Multiple imputation additionally reflects the uncertainty of the estimate rather than pretending the imputed value is known.

Whatever you choose, report it: how much was missing, on which variables, what mechanism you judged it to be and why, what treatment you applied, and the analysis sample size that resulted. Where the mechanism may be non-random, say so and treat it as a limitation, since no technique repairs it and disclosure is the only honest response.

Note: Missing data on the dependent variable are different from missing data on predictors. Cases with no outcome value contribute nothing to a regression and imputing the outcome imports assumptions into the thing you are trying to explain.
الرابع

القيم المفقودة

البيانات المفقودة ليست إزعاجًا يُزاح. فنمط ما هو مفقود معلومة، والمعالجة الصحيحة تعتمد كليًا على لماذا غابت القيم.

مفقود عشوائيًا تمامًامفقود عشوائيًامفقود لا عشوائيًاالفقدان لا صلة له بأي شيءالفقدان تفسّره متغيرات أخرىلديكالفقدان يعتمد على القيمةالمفقودة نفسهاإسقاط الحالات غير متحيّزالتعويض قد ينجح إن كانتالمتنبئات في البياناتلا تقنية تصلحه؛ ويجب الإفصاح عنهثلاث آليات؛ والثالثة وحدها قاتلة، والتعليل وحده يحددها

ابدأ بـالتكميم: نسبة المفقود لكل متغير ولكل حالة، وهل يتعنقد الفقدان. فمتغيرٌ مفقود في عشرين بالمئة من الحالات مشكلة؛ ومتغيرٌ مفقود في اثنين بالمئة ليس كذلك. والحالات التي تفقد قيمًا كثيرة يُفضَّل إسقاطها على إصلاحها عادةً.

ثم علّل عن الآلية، وهي لا تُحدَّد إحصائيًا ويجب أن يُحاجّ بها. فإن كان الدخل مفقودًا أكثر لدى الأعلى دخلًا، فالفقدان يعتمد على القيمة المفقودة نفسها وليس عشوائيًا بالمعنى الفني. وإن كانت كتلة بنود مفقودة لكل من أخذ المسح على جهاز محمول، فالفقدان تفسّره متغير لديك، ويمكن معالجته. وإن غابت القيم بسبب خطأ إرسال، فالفقدان لا صلة له بشيء.

والمعالجات تتبع الآلية. الحذف القائمي، أي إسقاط أي حالة بقيمة مفقودة على متغيرات التحليل، هو الافتراضي في معظم البرمجيات وهو غير متحيّز فقط حين تكون البيانات مفقودة عشوائيًا تمامًا. وكلفته الحقيقية حجم العيّنة: فبعشرة متغيرات كلٌّ مفقود بخمسة بالمئة، قد يزيل الحذف القائمي ثلث العيّنة. والحذف الزوجي يستخدم كل البيانات المتاحة لكل حساب وينتج مصفوفة ارتباط تقوم فيها خانات مختلفة على حالات مختلفة، وهذا يسبب مشكلات في التقنيات متعددة المتغيرات.

وإبدال المتوسط، أي استبدال متوسط المتغير بقيمة مفقودة، شائع في عمل الطلاب وغير موصى به: فهو يخفّض التباين صناعيًا ويُضعف العلاقات نفسها التي تختبرها. والتعويض الانحداري أو المتعدد يقدّر القيم المفقودة من المتغيرات الأخرى وهو المقاربة الحديثة المفضلة حين يفسّر الفقدانَ متغيراتٌ لديك. والتعويض المتعدد يعكس إضافةً إلى ذلك عدم يقين التقدير بدل التظاهر بأن القيمة المعوَّضة معلومة.

وأيًّا كان اختيارك، أبلغ به: كم كان المفقود، وعلى أي متغيرات، وأي آلية حكمت بأنها ولماذا، وأي معالجة طبّقت، وحجم عيّنة التحليل الناتج. وحيث قد تكون الآلية لا عشوائية، فقُل ذلك وعامله حدًّا، لأن لا تقنية تصلحه والإفصاح هو الاستجابة الصادقة الوحيدة.

ملاحظة: البيانات المفقودة على المتغير التابع غير المفقودة على المتنبئات. فالحالات بلا قيمة مخرَج لا تسهم بشيء في انحدار وتعويضُ المخرَج يُدخل افتراضات في الشيء الذي تحاول تفسيره.
Five

Recoding, Scoring, and Assumption Checks

This step builds the variables you will actually analyse, and it is where a small mistake propagates silently into every result.

Reverse-scored items must be reversed before any scale is computed. On a five-point scale the transformation is six minus the response. Forgetting this on one item of a five-item scale destroys the scale's reliability and, more dangerously, sometimes does not destroy it enough to be noticed. The check is to compute the correlation of each item with the scale total: an item correlating negatively has not been reversed.

Scale scores are usually the mean of the constituent items rather than the sum, because the mean stays on the original response metric and is interpretable. Decide in advance how many missing items a participant may have and still receive a score; a common rule is that a mean is computed when at least eighty percent of items are answered.

Derived variables should be built and checked one at a time. Categorical variables need dummy coding with a stated reference category. Interaction terms need their components centred if you intend to interpret the main effects. Any computed variable should be checked against a handful of cases by hand, because a formula error in a derived variable is invisible in the output and fatal to the conclusion.

Assumption checks come before running any test, not after an unwelcome result. Which assumptions apply depends on the technique, and four recur. Normality of the residuals matters for parametric tests, and is checked with a plot rather than only with a significance test, since tests of normality reject trivial departures in large samples. Homogeneity of variance across groups matters for comparisons of means. Linearity matters for correlation and regression, and a scatterplot reveals a curved relationship that a coefficient reports as weak. And independence of observations matters for almost everything, and is a design property rather than something you can test after the fact.

For regression, add multicollinearity, which is checked with variance inflation factors and matters because highly correlated predictors produce unstable coefficients that change sign between samples. If two predictors correlate above about 0.80, they are probably measuring one thing, and the analytical response is to combine them or drop one rather than to report both.

Where an assumption fails, three responses are available and all are legitimate if stated: transform the variable, use a technique that does not require the assumption, or proceed and report the violation as a limitation. What is not legitimate is checking the assumption, finding it violated, and reporting the test as though it held.

Bottom line: decide every preparation rule before looking at the outcomes, and never edit the raw file. Work in six steps: import and inspect, screen cases, handle missing values, recode and score, check assumptions, freeze and document. Screen with thresholds fixed in advance, and distinguish error outliers from ineligible ones from legitimate extremes, reporting results with and without the last kind. Reason about the missing data mechanism rather than defaulting to deletion, reverse items before computing scales, check derived variables by hand, and run assumption checks before the test rather than after the result.
الخامس

إعادة الترميز والتسجيل وفحوص الافتراضات

هذه الخطوة تبني المتغيرات التي ستحللها فعلًا، وهي حيث ينتشر خطأ صغير بصمت إلى كل نتيجة.

والبنود المعكوسة التسجيل يجب عكسها قبل حساب أي مقياس. وعلى مقياس خماسي يكون التحويل ستة ناقص الاستجابة. ونسيان هذا في بند واحد من مقياس خماسي يدمّر ثبات المقياس، والأخطر أنه أحيانًا لا يدمّره بما يكفي ليُلاحَظ. والفحص حسابُ ارتباط كل بند بمجموع المقياس: فالبند المرتبط سلبًا لم يُعكَس.

ودرجات المقاييس تكون عادةً متوسط البنود المكوِّنة لا مجموعها، لأن المتوسط يبقى على مقياس الاستجابة الأصلي وهو قابل للتفسير. وقرّر سلفًا كم بندًا مفقودًا يجوز لمشارك أن يفقده ويظل يتلقى درجة؛ والقاعدة الشائعة أن يُحسَب المتوسط حين تُجاب ثمانون بالمئة من البنود على الأقل.

والمتغيرات المشتقة ينبغي بناؤها وفحصها واحدًا واحدًا. فالمتغيرات الفئوية تحتاج ترميزًا صوريًا بفئة مرجعية مذكورة. وحدود التفاعل تحتاج توسيط مكوناتها إن كنت تعتزم تفسير الآثار الرئيسة. وأي متغير محسوب ينبغي فحصه مقابل حفنة حالات يدويًا، لأن خطأ صيغة في متغير مشتق غير مرئي في المخرَج وقاتل للخاتمة.

وفحوص الافتراضات تأتي قبل تشغيل أي اختبار لا بعد نتيجة غير مرحَّب بها. وأي الافتراضات ينطبق يعتمد على التقنية، وأربعة تتكرر. الاعتدالية للبواقي تهم للاختبارات المعلمية، وتُفحَص برسم لا باختبار دلالة وحده، لأن اختبارات الاعتدالية ترفض انحرافات تافهة في العيّنات الكبيرة. وتجانس التباين عبر المجموعات يهم لمقارنات المتوسطات. والخطية تهم للارتباط والانحدار، ورسمُ الانتشار يكشف علاقة منحنية يُبلغ عنها المعامل ضعيفةً. واستقلال الملاحظات يهم لكل شيء تقريبًا، وهو خاصية تصميم لا شيءٌ تستطيع اختباره بعد الوقوع.

وللانحدار، أضف التعدد الخطي، ويُفحَص بمعاملات تضخم التباين ويهم لأن المتنبئات عالية الارتباط تُنتج معاملات غير مستقرة تغيّر إشارتها بين العيّنات. فإن ارتبط متنبئان فوق نحو 0.80، فهما على الأرجح يقيسان شيئًا واحدًا، والاستجابة التحليلية دمجُهما أو إسقاط أحدهما لا الإبلاغ بكليهما.

وحيث يخفق افتراض، ثلاث استجابات متاحة وكلها مشروعة إن ذُكرت: حوّل المتغير، أو استخدم تقنية لا تتطلب الافتراض، أو امضِ وأبلغ بالانتهاك حدًّا. وما ليس مشروعًا فحصُ الافتراض ووجدانُه منتهَكًا والإبلاغُ بالاختبار وكأنه صحيح.

الخلاصة: قرّر كل قاعدة إعداد قبل النظر في المخرجات، ولا تحرر الملف الخام أبدًا. اعمل في ست خطوات: الاستيراد والفحص، وفرز الحالات، ومعالجة المفقود، وإعادة الترميز والتسجيل، وفحص الافتراضات، والتجميد والتوثيق. وافرز بعتبات مثبَّتة سلفًا، وميّز الشواذ الخطأ عن غير المؤهلة عن المتطرفات المشروعة، مُبلِغًا بالنتائج مع النوع الأخير وبدونه. وعلّل عن آلية الفقدان بدل التخلف إلى الحذف، واعكس البنود قبل حساب المقاييس، وافحص المتغيرات المشتقة يدويًا، وأجرِ فحوص الافتراضات قبل الاختبار لا بعد النتيجة.