A questionnaire converts a concept into numbers. Everything that happens afterwards, every statistic, every finding, every conclusion, inherits whatever the instrument did well or badly, and no amount of sophisticated analysis repairs a measure that did not capture the construct.
Two errors account for most weak instruments. The first is starting from items rather than from constructs: writing questions that seem relevant before stating precisely what each variable means. An item can only be judged against a definition, and without one there is no way to tell whether a question belongs in the instrument.
The second is inventing scales that already exist. Business research has validated measures for most common constructs, developed over years and tested on thousands of respondents. A scale you write yourself has no validity evidence, no comparability with other studies, and no defence when an examiner asks how you know it measures what you say. Borrowing is not laziness; it is the standard and expected practice, and inventing is what requires justification.
This framework covers defining constructs so items can be written against them, finding and adapting existing scales, the rules of item wording, the four levels of measurement and what each permits, questionnaire structure and layout, and piloting. It is my own synthesis, written in my own words and grounded in recognized scholarship.
الاستبانة تحوّل مفهومًا إلى أرقام. وكل ما يحدث بعدها، من إحصاءة ونتيجة وخاتمة، يرث ما أحسنته الأداة أو أساءته، ولا قدرَ من التحليل المتطور يُصلح مقياسًا لم يلتقط البناء.
وخطآن يفسّران معظم الأدوات الضعيفة. الأول البدء من البنود لا من البناءات: كتابةُ أسئلة تبدو ذات صلة قبل ذكر ما يعنيه كل متغير بالضبط. فالبند لا يمكن الحكم عليه إلا مقابل تعريف، وبدون تعريف لا سبيل لمعرفة هل ينتمي سؤالٌ إلى الأداة.
والثاني ابتكار مقاييس موجودة سلفًا. فلبحوث الأعمال مقاييس مصدّقة لمعظم البناءات الشائعة، طُوِّرت عبر سنوات واختُبرت على آلاف المستجيبين. والمقياس الذي تكتبه بنفسك بلا دليل صدق، ولا قابلية مقارنة بدراسات أخرى، ولا دفاع حين يسأل ممتحن كيف تعرف أنه يقيس ما تقول. والاستعارة ليست كسلًا؛ بل هي الممارسة المعيارية والمتوقّعة، والابتكار هو ما يتطلب تبريرًا.
ويغطي هذا الإطار تعريف البناءات لتُكتَب البنود مقابلها، وإيجاد المقاييس القائمة وتكييفها، وقواعد صياغة البنود، ومستويات القياس الأربعة وما يجيزه كلٌّ، وبنية الاستبانة وتخطيطها، والاختبار التجريبي. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
A construct is an abstract concept you cannot observe directly. Measuring it means specifying observable indicators, and specification has to come first.
The procedure has three steps and takes an afternoon. Write a conceptual definition for each variable, drawn from the literature and cited: job satisfaction is a pleasurable emotional state resulting from the appraisal of one's job. This definition governs everything that follows, and where the literature offers competing definitions you must choose one and say why, because a study that measures the construct under one definition and discusses it under another is incoherent.
Specify the dimensions. Many business constructs are multidimensional: organizational commitment is conventionally treated as affective, continuance, and normative, and a study measuring only one must say so rather than call the result commitment. If your construct has dimensions, decide whether you measure all of them and whether your analysis treats them separately or combines them.
Write the operational definition, which states how the construct will be observed: job satisfaction is measured as the mean of five items from a named scale, each rated on a five-point agreement format. This sentence is what connects the abstract concept to the numbers in your dataset, and it is the sentence a reader needs in order to judge whether your findings are about what you say they are about.
Two failures follow from skipping these steps. Construct drift: the definition in chapter two and the items in chapter three measure slightly different things, which nobody notices until the discussion contradicts the literature review. And the unmeasured dimension: a claim about a whole construct built on items covering one part of it, which is the most common validity problem in student instruments.
Test the definitions by writing, for each construct, one sentence describing someone who scores high and one describing someone who scores low. If those two descriptions are not clearly different, the construct is not yet defined well enough to measure, and no set of items will fix that.
البناء مفهومٌ مجرد لا تستطيع ملاحظته مباشرةً. وقياسه يعني تحديد مؤشرات ملحوظة، والتحديد يجب أن يأتي أولًا.
وللإجراء ثلاث خطوات ويستغرق بعد ظهيرة. اكتب تعريفًا مفاهيميًا لكل متغير، مأخوذًا من الأدبيات ومستشهَدًا به: «الرضا الوظيفي حالة انفعالية سارّة ناتجة عن تقييم المرء لوظيفته». ويحكم هذا التعريف كل ما يليه، وحيث تعرض الأدبيات تعريفات متنافسة عليك اختيار واحد وقول لماذا، لأن دراسةً تقيس البناء بتعريف وتناقشه بآخر دراسةٌ غير متماسكة.
حدّد الأبعاد. فكثير من بناءات الأعمال متعدد الأبعاد: فالالتزام التنظيمي يُعامَل عرفًا وجدانيًا واستمراريًا ومعياريًا، والدراسة التي تقيس واحدًا فقط يجب أن تقول ذلك لا أن تسمّي النتيجة «التزامًا». فإن كان لبنائك أبعاد، فقرّر هل تقيسها كلها وهل يعاملها تحليلك منفصلةً أم يجمعها.
واكتب التعريف الإجرائي، الذي يذكر كيف سيُلاحَظ البناء: «يُقاس الرضا الوظيفي متوسطَ خمسة بنود من مقياس مسمّى، كلٌّ مقدَّر بصيغة موافقة خماسية». وهذه الجملة هي ما يصل المفهوم المجرد بالأرقام في مجموعة بياناتك، وهي الجملة التي يحتاجها القارئ للحكم هل نتائجك عمّا تقول إنها عنه.
وإخفاقان يتبعان تخطّي هذه الخطوات. انجراف البناء: فالتعريف في الفصل الثاني والبنود في الفصل الثالث تقيس أشياء مختلفة قليلًا، ولا يلاحظ ذلك أحد حتى تناقض المناقشةُ مراجعةَ الأدبيات. والبعد غير المقيس: ادّعاءٌ عن بناء كامل مبنيٌّ على بنود تغطي جزءًا منه، وهذه أشيع مشكلة صدق في أدوات الطلاب.
واختبر التعريفات بكتابة جملة لكل بناء تصف من يسجّل درجةً عالية وجملة تصف من يسجّل درجة منخفضة. فإن لم يكن الوصفان مختلفين بوضوح، فالبناء لم يُعرَّف بعد بما يكفي للقياس، ولن تُصلح ذلك أي مجموعة بنود.
Established scales carry validity evidence you cannot generate yourself in a doctorate. Using them is the default, and the work is finding the right one and adapting it honestly.
Find scales by looking in the papers you already have. Empirical studies name their instruments in the methods section, and the papers most similar to yours have already solved the measurement problem for your construct. Where several scales exist, choose on three criteria: how closely the scale's definition matches yours, how much validity evidence it carries, and how long it is, since a fifty-item measure of one construct is unusable in a multi-construct survey.
Record the provenance for each scale you use: the original source, the number of items, the response format, the reported reliability in prior studies, and any subsequent validation in a context like yours. This is what you report in the methodology, and it is what converts the phrase a validated scale into an actual justification.
Adaptation is normal and must be disclosed. Three kinds occur. Contextual wording changes the referent: my supervisor becomes my department head, or this company becomes this organization. This is minor and usually safe, but it should still be reported. Item reduction shortens a long scale, which is common and risky, because a short form's reliability is not the long form's reliability and must be checked in your own data. Translation is the largest adaptation and requires a specific procedure.
The standard translation procedure is forward and back translation. One bilingual translator renders the scale into the target language; a second, working independently and without seeing the original, translates it back; the two English versions are compared and discrepancies resolved by discussion; the resulting version is reviewed by subject experts and then piloted with respondents in the target language. Report this process, because a scale translated informally has unknown properties and the reliability reported in the original language does not transfer.
Where no suitable scale exists you must develop one, and the honest response is to treat that as a component of the study rather than a paragraph. Development requires generating items from the definition and from qualitative data, expert review for content validity, a pilot with enough respondents to run a factor analysis, and reporting of the resulting structure and reliability. If that is beyond the scope of your project, the correct response is usually to redefine the construct to one that has been measured rather than to invent items and hope.
المقاييس الراسخة تحمل أدلة صدق لا تستطيع توليدها بنفسك في دكتوراه. واستخدامها هو الافتراضي، والعمل هو إيجاد المناسب وتكييفه بأمانة.
جِد المقاييس بالنظر في الأوراق التي تملكها سلفًا. فالدراسات التجريبية تسمّي أدواتها في قسم المناهج، والأوراق الأشبه بدراستك حلّت مشكلة القياس لبنائك سلفًا. وحيث توجد مقاييس عدة، اختر بثلاثة معايير: مدى مطابقة تعريف المقياس لتعريفك، وقدر أدلة الصدق التي يحملها، وطوله، لأن مقياسًا من خمسين بندًا لبناء واحد غير صالح للاستعمال في مسح متعدد البناءات.
سجّل المصدر لكل مقياس تستخدمه: المصدر الأصلي، وعدد البنود، وصيغة الاستجابة، والثبات المُبلَّغ عنه في دراسات سابقة، وأي تصديق لاحق في سياق كسياقك. وهذا ما تُبلغ به في المنهجية، وهو ما يحوّل عبارة «مقياس مصدَّق» إلى تبرير فعلي.
والتكييف طبيعي ويجب الإفصاح عنه. وثلاثة أنواع تقع. صياغة السياق تغيّر المرجع: «مشرفي» تصير «رئيس قسمي»، أو «هذه الشركة» تصير «هذه المنظمة». وهذا طفيف وآمن عادةً، لكنه ما زال ينبغي الإبلاغ عنه. وتقليل البنود يختصر مقياسًا طويلًا، وهذا شائع وخطر، لأن ثبات الصيغة القصيرة ليس ثبات الطويلة ويجب فحصه في بياناتك أنت. والترجمة أكبر التكييفات وتتطلب إجراءً محدَّدًا.
والإجراء المعياري للترجمة هو الترجمة الأمامية والعكسية. فمترجمٌ ثنائي اللغة ينقل المقياس إلى اللغة الهدف؛ وثانٍ، يعمل مستقلًا دون رؤية الأصل، يترجمه عكسيًا؛ وتُقارَن النسختان الإنجليزيتان وتُحَل الفروق بالنقاش؛ وتُراجَع النسخة الناتجة من خبراء الموضوع ثم تُجرَّب مع مستجيبين باللغة الهدف. أبلغ عن هذه العملية، لأن مقياسًا مترجَمًا بلا رسمية له خصائص مجهولة والثبات المُبلَّغ عنه في اللغة الأصل لا ينتقل.
وحيث لا يوجد مقياس مناسب عليك تطوير واحد، والاستجابة الصادقة معاملةُ ذلك مكوّنًا من الدراسة لا فقرةً. فالتطوير يتطلب توليد بنود من التعريف ومن بيانات نوعية، ومراجعةَ خبراء لصدق المحتوى، واختبارًا تجريبيًا بمستجيبين يكفون لتشغيل تحليل عاملي، وإبلاغًا بالبنية الناتجة والثبات. فإن كان ذلك خارج نطاق مشروعك، فالاستجابة الصحيحة عادةً إعادةُ تعريف البناء إلى واحد قيس سلفًا لا ابتكارُ بنود والأمل.
Where you do write items, the rules are few, well established, and violated constantly. Each rule exists because breaking it produces a specific measurable error.
| Rule | Bad | Better |
|---|---|---|
| One idea per item | My manager is supportive and fair | My manager is supportive |
| No leading wording | Do you agree the new system is an improvement? | How would you rate the new system? |
| Plain language | Does your firm leverage synergistic capabilities? | Do teams in your firm work together on projects? |
| No hidden assumption | How often does your manager give feedback? | Does your manager give feedback? If yes, how often? |
| Concrete referent | Are you satisfied with support? | Are you satisfied with the support from your direct manager? |
| No negation stacking | I do not think the policy is unhelpful | The policy is helpful |
The double-barrelled item, first in the table, is the most common and the most damaging, because a respondent who agrees with one half and disagrees with the other has no valid answer and will guess. Scan every item for the word and, and split anything that contains two claims.
Response formats should be consistent across the instrument wherever possible, because switching formats increases errors and slows completion. The five-point agreement format is standard and adequate for most purposes; seven points add discrimination for respondents who use scales carefully and add noise for those who do not. Even-numbered scales remove the midpoint and force a direction, which is appropriate when you believe a neutral response is usually evasion and inappropriate when neutrality is a genuine position.
Include a not applicable option where an item may not apply, since respondents without one will either skip the item, producing missing data, or answer anyway, producing wrong data. The second is worse and is invisible in the dataset.
Reverse-worded items are included in many published scales to disrupt automatic responding. Keep them if the source scale has them, remember to reverse the scoring before analysis, and check in your pilot that they behave as expected, because reverse items sometimes load as a separate factor, which is a known artefact rather than a finding.
Two structural rules complete the instrument. Order: start with easy, non-threatening items to build momentum, place the most important constructs early while attention is high, group items by topic, and put demographics at the end, where they feel less like screening. Length: aim for under fifteen minutes, since completion falls sharply beyond that, and state the estimated time in the invitation because respondents who know the cost are likelier to start and much likelier to finish.
وحيث تكتب بنودًا، فالقواعد قليلة وراسخة وتُنتهَك باستمرار. وكل قاعدة موجودة لأن كسرها يُنتج خطأً محدَّدًا قابلًا للقياس.
| القاعدة | سيئ | أفضل |
|---|---|---|
| فكرة واحدة لكل بند | مديري داعم وعادل | مديري داعم |
| لا صياغة موجِّهة | هل توافق أن النظام الجديد تحسين؟ | كيف تقيّم النظام الجديد؟ |
| لغة بسيطة | هل تستثمر منشأتك القدرات التآزرية؟ | هل تعمل الفرق في منشأتك معًا على المشاريع؟ |
| لا افتراض مخفي | كم مرة يعطيك مديرك تغذية راجعة؟ | هل يعطيك مديرك تغذية راجعة؟ وإن نعم، كم مرة؟ |
| مرجع محسوس | هل أنت راضٍ عن الدعم؟ | هل أنت راضٍ عن الدعم من مديرك المباشر؟ |
| لا تراكم نفي | لا أظن أن السياسة غير مفيدة | السياسة مفيدة |
والبند المزدوج، الأول في الجدول، أشيعها وأشدها ضررًا، لأن المستجيب الذي يوافق على نصف ويخالف الآخر لا إجابة صالحة لديه وسيخمّن. امسح كل بند بحثًا عن حرف العطف، وقسّم كل ما يحوي ادّعاءين.
وصيغ الاستجابة ينبغي أن تكون متسقة عبر الأداة حيثما أمكن، لأن تبديل الصيغ يزيد الأخطاء ويبطئ الإكمال. وصيغة الموافقة الخماسية معيارية وكافية لمعظم الأغراض؛ والسبع نقاط تضيف تمييزًا للمستجيبين الذين يستخدمون المقاييس بعناية وتضيف ضوضاءً لمن لا يفعلون. والمقاييس الزوجية تزيل نقطة الوسط وتفرض اتجاهًا، وهذا يلائم حين تعتقد أن الاستجابة المحايدة تهرّبٌ عادةً ولا يلائم حين يكون الحياد موقفًا حقيقيًا.
وأدرج خيار لا ينطبق حيث قد لا ينطبق بند، لأن المستجيبين بدونه إما سيتخطون البند فتنتج بيانات مفقودة، أو سيجيبون على أي حال فتنتج بيانات خاطئة. والثاني أسوأ وغير مرئي في مجموعة البيانات.
والبنود المعكوسة الصياغة مُدرَجة في مقاييس منشورة كثيرة لتعطيل الاستجابة الآلية. أبقِها إن كانت في المقياس المصدر، وتذكّر عكس التسجيل قبل التحليل، وافحص في تجربتك أنها تسلك كما هو متوقَّع، لأن البنود المعكوسة تُحمَّل أحيانًا عاملًا منفصلًا، وهذا أثرٌ صناعي معروف لا نتيجة.
وقاعدتان بنيويتان تكملان الأداة. الترتيب: ابدأ ببنود سهلة غير مهدِّدة لبناء الزخم، وضع أهم البناءات مبكرًا حين يكون الانتباه عاليًا، واجمع البنود بالموضوع، وضع البيانات الديموغرافية في النهاية حيث تبدو أقل شبهًا بالفرز. والطول: استهدف أقل من خمس عشرة دقيقة، فالإكمال يهبط بحدة بعد ذلك، واذكر الوقت المقدَّر في الدعوة لأن المستجيبين الذين يعرفون الكلفة أميل للبدء وأميل بكثير للإنهاء.
The level of measurement of each variable determines which statistics are legitimate on it. Getting this wrong invalidates an analysis in a way no reviewer overlooks.
Nominal variables are categories with no order, and the only legitimate statistics are counts, proportions, and the mode. Coding sectors as 1 to 5 does not make them numbers; a mean sector of 3.2 is meaningless. Ordinal variables are ordered but with unequal gaps, so the median and rank-based tests are appropriate while the mean is strictly not.
The practical difficulty concerns rating scales. A single five-point agreement item is ordinal: the distance between agree and strongly agree is not known to equal the distance between neutral and agree. In practice, business research routinely treats the mean of several such items as interval, which is a defensible convention because averaging multiple indicators approximates an underlying continuum. State that you are doing this, treat single items as ordinal, and use the mean of a multi-item scale where you need interval-level techniques.
Interval and ratio variables support the full range of parametric statistics. The difference between them matters only for ratios of values: it is meaningful to say one employee has twice the tenure of another, because tenure has a true zero, and not meaningful to say one has twice the satisfaction.
Two practical consequences follow. Collect at the highest level available, since you can always collapse later but never recover detail you did not collect. Ask for exact age or tenure rather than bands unless there is a privacy reason for bands, and if you use bands, say why. And state each variable's level in your instrument table, because that column is what justifies the statistics you run in chapter four.
Piloting is the final step and the one most often skipped. Run the instrument on ten to thirty people resembling the target sample. Time completion, since the true duration is invariably longer than the estimate. Ask respondents to think aloud, which reveals items understood differently from your intention faster than any other technique. Check the data for items with no variance, since an item everyone answers identically measures nothing. And where the pilot is large enough, compute reliability for each multi-item scale.
Report the pilot in the methodology: how many people, who they were, what was tested, and what changed as a result. A pilot that changed nothing is a pilot that was not really run, and reporting the two or three items you rewrote is stronger evidence of care than any assertion that the instrument was validated.
مستوى قياس كل متغير يحدد أي الإحصاءات مشروعة عليه. والخطأ في هذا يُبطل تحليلًا بطريقة لا يغفلها محكّم.
والمتغيرات الاسمية فئات بلا ترتيب، والإحصاءات المشروعة الوحيدة الأعدادُ والنسبُ والمنوال. فترميز القطاعات من 1 إلى 5 لا يجعلها أرقامًا؛ ومتوسط قطاع 3.2 بلا معنى. والمتغيرات الرتبية مرتبة لكن بفجوات غير متساوية، فالوسيط والاختبارات القائمة على الرتب تلائم بينما المتوسط لا يلائم بدقة.
والصعوبة العملية تخص مقاييس التقدير. فبندٌ خماسي مفرد للموافقة رتبي: فالمسافة بين «أوافق» و«أوافق بشدة» غير معلومة المساواة للمسافة بين «محايد» و«أوافق». وعمليًا، تعامل بحوث الأعمال روتينيًا متوسط عدة بنود كهذه فتريًا، وهذا عرفٌ قابل للدفاع لأن متوسط مؤشرات متعددة يقارب متصلًا كامنًا. اذكر أنك تفعل ذلك، وعامل البنود المفردة رتبيًا، واستخدم متوسط مقياس متعدد البنود حيث تحتاج تقنيات فترية.
والفترية والنسبية تسندان المدى الكامل للإحصاءات المعلمية. والفرق بينهما لا يهم إلا لنسب القيم: فمن المعنيّ قول إن موظفًا لديه ضعف مدة خدمة آخر، لأن مدة الخدمة لها صفر حقيقي، وليس معنيًّا قول إن لديه ضعف الرضا.
ونتيجتان عمليتان تتبعان. اجمع عند أعلى مستوى متاح، فتستطيع دائمًا الطي لاحقًا ولا تستطيع أبدًا استعادة تفصيل لم تجمعه. اسأل عن العمر أو مدة الخدمة بالضبط لا بفئات ما لم يكن ثمة سبب خصوصية للفئات، وإن استخدمت فئات فقُل لماذا. واذكر مستوى كل متغير في جدول أداتك، لأن ذلك العمود هو ما يبرّر الإحصاءات التي تشغّلها في الفصل الرابع.
والاختبار التجريبي الخطوة الأخيرة والأكثر تخطّيًا. شغّل الأداة على عشرة إلى ثلاثين شخصًا يشبهون العيّنة المستهدفة. وقِس زمن الإكمال، فالمدة الحقيقية أطول من التقدير دائمًا. واطلب من المستجيبين التفكير بصوت عالٍ، وهذا يكشف البنود المفهومة على غير قصدك أسرع من أي تقنية أخرى. وافحص البيانات بحثًا عن بنود بلا تباين، فالبند الذي يجيبه الجميع بالطريقة نفسها لا يقيس شيئًا. وحيث تكون التجربة كبيرة بما يكفي، احسب الثبات لكل مقياس متعدد البنود.
وأبلغ عن التجربة في المنهجية: كم شخصًا، ومن كانوا، وما اختُبر، وما تغيّر نتيجةً لذلك. فتجربةٌ لم تغيّر شيئًا تجربةٌ لم تُجرَ فعلًا، والإبلاغ بالبندين أو الثلاثة التي أعدت كتابتها دليلُ عنايةٍ أقوى من أي جزم بأن الأداة «صُدِّقت».