Software does not analyse data. It executes the analysis you specify, which means the tool never rescues a design and never substitutes for understanding what a technique does. The most common software problem in doctoral work is not choosing the wrong tool; it is expecting the tool to make a decision that belongs to the researcher.
That said, the choice matters practically. A tool you cannot use consumes months. A tool that cannot do what your design requires forces the design to change. And a workflow that cannot be rerun means that discovering an error late costs weeks rather than minutes, which is the single largest hidden cost in quantitative doctoral work.
Three principles govern the choice. Choose for the analysis, not for prestige. A survey study needing descriptives, correlations, and regression is served completely by a standard statistical package, and learning a programming language for it is a hobby rather than a requirement. Choose what you can get support for, meaning what your supervisor and department use, because a tool nobody around you knows means every problem is solved alone. And choose once, because migrating an analysis between tools mid-project is expensive and produces two versions of everything.
This framework covers the five categories of tool and what each is for, the specific role of qualitative software and what it does not do, reproducibility as a workflow rather than a tool, and what to report about software in a methodology chapter. It is my own synthesis, written in my own words and grounded in recognized scholarship.
البرمجيات لا تحلل البيانات. إنها تنفّذ التحليل الذي تحدده، ما يعني أن الأداة لا تنقذ تصميمًا أبدًا ولا تحل محل فهم ما تفعله تقنية. وأشيع مشكلة برمجية في عمل الدكتوراه ليست اختيار الأداة الخطأ؛ بل توقّعُ أن تتخذ الأداةُ قرارًا يخص الباحث.
ومع ذلك، الاختيار يهم عمليًا. فأداةٌ لا تستطيع استخدامها تستهلك أشهرًا. وأداةٌ لا تستطيع فعل ما يتطلبه تصميمك تُجبر التصميم على التغير. وسير عملٍ لا يمكن إعادة تشغيله يعني أن اكتشاف خطأ متأخرًا يكلّف أسابيع لا دقائق، وهذه أكبر كلفة خفية مفردة في عمل الدكتوراه الكمي.
وثلاثة مبادئ تحكم الاختيار. اختر للتحليل لا للهيبة. فدراسة مسح تحتاج وصفيات وارتباطات وانحدارًا تخدمها حزمةٌ إحصائية معيارية تمامًا، وتعلّمُ لغة برمجة لها هوايةٌ لا مطلب. واختر ما تستطيع الحصول على دعم له، أي ما يستخدمه مشرفك وقسمك، لأن أداةً لا يعرفها أحد حولك تعني حل كل مشكلة وحدك. واختر مرة واحدة، لأن ترحيل تحليل بين أدوات في منتصف المشروع مكلف ويُنتج نسختين من كل شيء.
ويغطي هذا الإطار فئات الأدوات الخمس ووظيفة كل واحدة، والدور المحدد للبرمجيات النوعية وما لا تفعله، وقابلية إعادة الإنتاج بوصفها سير عمل لا أداة، وما يُبلَّغ عنه من البرمجيات في فصل منهجية. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
Tools divide into five kinds by what they are built to do. Most doctoral projects use two or three together, and knowing the boundary between them prevents using one for a job it does badly.
Spreadsheets are excellent for data entry, cleaning, simple description, and charts, and they are where most business data already lives. Their strengths are ubiquity and immediacy. Their weaknesses matter for analysis: statistical functions are limited and some are implemented in ways that differ from statistical convention, formulas are invisible unless you look at each cell, and there is no record of what you did. A spreadsheet is the right tool for preparing data and the wrong tool for inferential analysis in a thesis.
Statistical packages are built for the standard tests and present them through menus with output designed to be read. They handle everything a typical business survey study needs: descriptives, reliability, correlation, group comparisons, regression, factor analysis, and more. Their advantage is that the analysis is close to the textbook and the output is easy to report. Their disadvantage is that clicking through menus leaves no record, which is why every serious package also offers a syntax or command file and why you should use it.
Programming languages do anything, and everything they do is written down by construction. The cost is a learning curve measured in months rather than weeks. They are the right choice when your analysis needs a technique the packages do not offer, when you are working with large or awkwardly shaped data, or when reproducibility is itself part of the contribution. They are the wrong choice when a package would do the job and the language would consume a term.
Qualitative software manages coding, retrieval, and querying of text, and is discussed in the next section because what it does is widely misunderstood. Reference managers run alongside everything else and are covered elsewhere in this track.
A practical recommendation for most business doctorates: use a spreadsheet for entry and cleaning, a statistical package driven by a saved syntax file for the analysis, qualitative software if you have more than about fifteen interviews, and a reference manager throughout. That combination is learnable in a month, supported everywhere, and adequate for the large majority of designs.
تنقسم الأدوات خمسة أنواع بحسب ما بُنيت لفعله. ومعظم مشاريع الدكتوراه تستخدم اثنتين أو ثلاثًا معًا، ومعرفة الحد بينها تمنع استخدام واحدة لعمل تؤديه سيئًا.
والجداول الحسابية ممتازة لإدخال البيانات وتنظيفها والوصف البسيط والرسوم، وهي حيث تسكن معظم بيانات الأعمال سلفًا. وقوّتها الانتشار والفورية. وضعفها يهم للتحليل: فالدوال الإحصائية محدودة وبعضها مُنفَّذ بطرق تختلف عن العرف الإحصائي، والصيغ غير مرئية ما لم تنظر في كل خانة، ولا سجل لما فعلت. فالجدول الحسابي الأداة الصحيحة لإعداد البيانات والأداة الخطأ للتحليل الاستدلالي في أطروحة.
والحزم الإحصائية مبنيّة للاختبارات المعيارية وتعرضها عبر قوائم بمخرَج مصمَّم ليُقرأ. وهي تعالج كل ما تحتاجه دراسة مسح أعمال نموذجية: الوصفيات والثبات والارتباط ومقارنات المجموعات والانحدار والتحليل العاملي وأكثر. وميزتها أن التحليل قريب من الكتاب المدرسي والمخرَج سهل الإبلاغ. وعيبها أن النقر عبر القوائم لا يترك سجلًا، ولهذا تعرض كل حزمة جادة أيضًا ملف صياغة أو أوامر ولهذا ينبغي أن تستخدمه.
ولغات البرمجة تفعل أي شيء، وكل ما تفعله مكتوب بالبناء. والكلفة منحنى تعلّم يُقاس بالأشهر لا بالأسابيع. وهي الخيار الصحيح حين يحتاج تحليلك تقنيةً لا تعرضها الحزم، أو حين تعمل ببيانات كبيرة أو غريبة الشكل، أو حين تكون قابلية إعادة الإنتاج نفسها جزءًا من الإسهام. وهي الخيار الخطأ حين تؤدي حزمةٌ العمل وتستهلك اللغةُ فصلًا دراسيًا.
والبرمجيات النوعية تدير ترميز النصوص واسترجاعها والاستعلام عنها، وتُطرَح في القسم التالي لأن ما تفعله يُساء فهمه على نطاق واسع. ومديرو المراجع يعملون إلى جانب كل ما عداهم ويُغطَّون في موضع آخر من هذا المسار.
وتوصية عملية لمعظم دكتوراه الأعمال: استخدم جدولًا حسابيًا للإدخال والتنظيف، وحزمةً إحصائية يقودها ملف صياغة محفوظ للتحليل، وبرمجيةً نوعية إن كان لديك أكثر من نحو خمس عشرة مقابلة، ومدير مراجع طوال الوقت. وذلك المزيج قابل للتعلّم في شهر، ومدعوم في كل مكان، وكافٍ للغالبية العظمى من التصاميم.
Qualitative analysis software is widely misunderstood in both directions: some students expect it to analyse for them, and others avoid it believing it mechanises interpretation. Both are wrong.
| What it does | What it does not do |
|---|---|
| Stores transcripts, documents, images, and audio in one project | Decide what is interesting |
| Lets you attach codes to segments and retrieve everything under a code | Generate the codes for you |
| Shows which codes co-occur and where | Tell you what the co-occurrence means |
| Compares coding across participant groups | Establish that a difference is meaningful |
| Records who coded what and when | Make the coding consistent |
| Holds memos linked to codes and extracts | Write the analysis |
The value is in the middle column of that list: retrieval and comparison at scale. With twenty interviews and eighty codes, the question of what everyone said about autonomy is answerable in two seconds by software and takes an hour by hand. The question of whether managers and staff coded differently on the same theme is answerable at all only if the coding is in a system.
Two dangers are worth naming. The first is the coding trap, where the ease of coding leads to producing four hundred codes and never moving to the interpretive phases, because the software makes phase two feel productive and offers no equivalent support for phases three to five. The second is false quantification, where the software's frequency counts invite reporting that a code appeared thirty-seven times, which is a statement about your coding rather than about the world.
The threshold for using it is roughly fifteen interviews. Below that, coding in a word processor or on paper is entirely workable and many experienced researchers prefer it because it keeps them closer to the text. Above it, manual retrieval becomes the bottleneck and the software pays for itself.
Whichever you use, report it correctly. Say the software was used to organise and retrieve coded data, not that it was used to analyse the data, because the second claim attributes your interpretive work to a tool. The conventional sentence is that coding was conducted by the researcher and managed using the named software.
برمجيات التحليل النوعي يُساء فهمها في الاتجاهين: فبعض الطلاب يتوقعون أن تحلل نيابةً عنهم، وآخرون يتجنبونها اعتقادًا بأنها تُميكِن التفسير. وكلاهما خطأ.
| ما تفعله | ما لا تفعله |
|---|---|
| تخزّن التفريغات والوثائق والصور والصوت في مشروع واحد | تقرر ما هو مثير |
| تتيح إلحاق رموز بمقاطع واسترجاع كل ما تحت رمز | تولّد الرموز لك |
| تُظهر أي الرموز تتزامن وأين | تخبرك ما يعنيه التزامن |
| تقارن الترميز عبر فئات المشاركين | تثبت أن فرقًا ذو معنى |
| تسجّل من رمّز ماذا ومتى | تجعل الترميز متسقًا |
| تحفظ مذكرات مرتبطة بالرموز والمقتطفات | تكتب التحليل |
والقيمة في العمود الأول من تلك القائمة: الاسترجاع والمقارنة بمقياس واسع. فبعشرين مقابلة وثمانين رمزًا، يكون سؤال ما قاله الجميع عن الاستقلالية قابلًا للإجابة في ثانيتين بالبرمجية ويستغرق ساعة يدويًا. وسؤالُ هل رمّز المديرون والموظفون بشكل مختلف على الموضوع نفسه قابلٌ للإجابة أصلًا فقط إن كان الترميز في نظام.
وخطران يستحقان التسمية. الأول فخ الترميز، حيث تقود سهولةُ الترميز إلى إنتاج أربعمئة رمز وعدم الانتقال قط إلى المراحل التفسيرية، لأن البرمجية تجعل المرحلة الثانية تبدو منتجة ولا تقدّم سندًا مكافئًا للمراحل الثالثة إلى الخامسة. والثاني التكميم الزائف، حيث تدعو أعداد التكرار في البرمجية إلى الإبلاغ بأن رمزًا ظهر سبعًا وثلاثين مرة، وهذه عبارة عن ترميزك لا عن العالم.
وعتبة استخدامها نحو خمس عشرة مقابلة. فتحتها، يكون الترميز في معالج نصوص أو على ورق عمليًا تمامًا ويفضّله كثير من الباحثين الخبيرين لأنه يبقيهم أقرب إلى النص. وفوقها، يصير الاسترجاع اليدوي عنق الزجاجة وتسدّد البرمجية ثمنها.
وأيًّا كانت التي تستخدم، أبلغ بها صحيحًا. قُل إن البرمجية استُخدمت لتنظيم البيانات المرمَّزة واسترجاعها، لا إنها استُخدمت لتحليل البيانات، لأن الادّعاء الثاني ينسب عملك التفسيري إلى أداة. والجملة المتعارفة أن الترميز أجراه الباحث وأُديرَ باستخدام البرمجية المسمّاة.
Reproducibility is not a property of a tool. It is a way of arranging your work such that every number in the thesis can be regenerated from the raw data by running a defined sequence.
The practical test is a question you should be able to answer yes to: if your analysis file were deleted, could you rebuild it exactly from the raw data and your scripts? If the answer is no, some step exists only in your memory, and that step is where an unfindable error will eventually be found.
Four practices produce it. Keep raw data untouched, read-only, in a separate folder, and never edit it by hand even to fix an obvious typo; fix it in the script so the fix is recorded. Do every transformation in code, meaning a syntax file, a command file, or a script, so the path from raw to analysis file is written down. Regenerate output rather than saving it, so that a change upstream propagates instead of leaving one stale table in the document. And version everything, with dated filenames at minimum, so you can identify which version produced which table.
A simple folder structure supports all four: one folder for raw data, one for scripts, one for generated output, one for the document, and one for documentation including the codebook and the analysis log. This costs ten minutes to set up and it is the difference between a correction taking an hour and taking a week.
Two additional habits matter over a multi-year project. Back up in three places, one of them off-site, and test the restore at least once, because an untested backup is a belief rather than a backup. And write a short read-me in the project folder saying what each folder contains and in what order the scripts run, addressed to yourself in two years, who will remember none of it.
The payoff appears at the worst possible moment. Late in the project you will discover a coding error, a supervisor will ask for the analysis rerun on a subsample, or an examiner will request a robustness check. With a reproducible workflow each of those is an afternoon. Without one, each is a week of reconstructing what you did, with no certainty that the reconstruction matches.
قابلية إعادة الإنتاج ليست خاصية أداة. إنها طريقةُ ترتيب عملك بحيث يمكن إعادة توليد كل رقم في الأطروحة من البيانات الخام بتشغيل تسلسل محدَّد.
والاختبار العملي سؤالٌ ينبغي أن تستطيع إجابته بنعم: لو حُذف ملف تحليلك، أتستطيع إعادة بنائه بالضبط من البيانات الخام ونصوصك؟ فإن كانت الإجابة لا، فثمة خطوة توجد في ذاكرتك وحدها، وتلك الخطوة حيث سيُوجَد في النهاية خطأ لا يمكن إيجاده.
وأربع ممارسات تُنتجها. أبقِ البيانات الخام دون مساس، للقراءة فقط، في مجلد منفصل، ولا تحررها يدويًا حتى لإصلاح خطأ مطبعي بيّن؛ أصلحه في النص ليُسجَّل الإصلاح. وأجرِ كل تحويل بالشفرة، أي ملف صياغة أو ملف أوامر أو نص برمجي، ليُكتَب المسار من الخام إلى ملف التحليل. وأعد توليد المخرَج بدل حفظه، ليتنشر تغييرٌ في المنبع بدل ترك جدول واحد قديم في المستند. وأصدر كل شيء، بأسماء ملفات مؤرَّخة كحد أدنى، لتستطيع تحديد أي إصدار أنتج أي جدول.
وبنية مجلدات بسيطة تسند الأربعة: مجلد للبيانات الخام، وآخر للنصوص، وآخر للمخرَج المولَّد، وآخر للمستند، وآخر للتوثيق بما فيه دليل الترميز وسجل التحليل. وهذا يكلّف عشر دقائق للإعداد وهو الفرق بين تصحيح يستغرق ساعة وآخر يستغرق أسبوعًا.
وعادتان إضافيتان تهمان في مشروع متعدد السنوات. احتفظ بنسخ في ثلاثة مواضع، واحد منها خارج الموقع، واختبر الاسترجاع مرة على الأقل، لأن نسخةً احتياطية غير مختبَرة اعتقادٌ لا نسخة. واكتب ملف تعريف قصيرًا في مجلد المشروع يقول ما يحويه كل مجلد وبأي ترتيب تعمل النصوص، مُوجَّهًا إلى نفسك بعد سنتين، وهو لن يتذكر شيئًا منه.
والعائد يظهر في أسوأ لحظة ممكنة. ففي أواخر المشروع ستكتشف خطأ ترميز، أو سيطلب مشرف إعادة تشغيل التحليل على عيّنة فرعية، أو سيطلب ممتحن فحص متانة. ومع سير عمل قابل لإعادة الإنتاج يكون كلٌّ من ذلك بعد ظهيرة. وبدونه، يكون كلٌّ أسبوعًا من إعادة بناء ما فعلت، بلا يقين أن إعادة البناء تطابق.
Software reporting is short and is checked. Four items belong in the methodology and each takes a clause.
Name and version. Statistical packages change their defaults and occasionally their algorithms between versions, so a result may not reproduce exactly under a different one. Report the version number, not just the product name.
What each tool was used for. A study using three tools should say which did what: data preparation in one, inferential analysis in another, qualitative coding in a third. This takes one sentence and prevents the ambiguity of a methods section that lists software without connecting it to steps.
Non-default settings. Where you changed an option that affects results, report it: the rotation method in a factor analysis, the estimation method in a model, the treatment of missing values, the tie-breaking rule in a rank test. Defaults need not be reported; departures from them must be, since a reader assuming the default cannot reproduce your number.
Any add-on or macro. Extensions used for specific procedures should be named with their version, because they are separate software and are frequently the part a reader cannot obtain.
Three things not to do. Do not attribute analysis to software: the data were analysed using the named package is acceptable shorthand, but qualitative software analysed the interviews is not, because it did not. Do not report software instead of technique: naming a package tells a reader nothing about which test you ran, and the technique is what matters. And do not paste raw software output into the thesis, since output is designed for the analyst and a results chapter is written for the reader; extract the numbers you need into a table you designed.
One larger point closes this. Software competence is worth acquiring and is not the skill being assessed. An examiner does not ask which package you used; they ask why you chose that technique, whether its assumptions held, what the effect size means, and how the finding bears on your question. A thesis can be built with modest tools and excellent judgement, and cannot be rescued by sophisticated tools and poor judgement. Choose the tool that lets you spend your attention on the analysis rather than on the tool.
الإبلاغ بالبرمجيات قصير ويُفحَص. وأربعة بنود مكانها المنهجية وكلٌّ يستغرق عبارة.
الاسم والإصدار. فالحزم الإحصائية تغيّر افتراضاتها وأحيانًا خوارزمياتها بين الإصدارات، فقد لا تُعاد نتيجةٌ بالضبط تحت إصدار آخر. أبلغ برقم الإصدار لا باسم المنتج فحسب.
ما استُخدمت له كل أداة. فدراسةٌ تستخدم ثلاث أدوات ينبغي أن تقول أيها فعل ماذا: إعداد البيانات في واحدة، والتحليل الاستدلالي في أخرى، والترميز النوعي في ثالثة. وهذا يستغرق جملةً ويمنع التباس قسم مناهج يسرد برمجيات دون ربطها بالخطوات.
الإعدادات غير الافتراضية. فحيث غيّرت خيارًا يؤثر في النتائج، أبلغ به: منهج التدوير في تحليل عاملي، ومنهج التقدير في نموذج، ومعالجة القيم المفقودة، وقاعدة كسر التعادل في اختبار رتبي. والافتراضات لا يلزم الإبلاغ بها؛ والانحرافات عنها يلزم، لأن قارئًا يفترض الافتراضي لا يستطيع إعادة إنتاج رقمك.
أي إضافة أو ماكرو. فالامتدادات المستخدَمة لإجراءات بعينها ينبغي تسميتها بإصدارها، لأنها برمجيات منفصلة وهي غالبًا الجزء الذي لا يستطيع القارئ الحصول عليه.
وثلاثة أشياء لا تُفعَل. لا تنسب التحليل إلى البرمجية: فـ«حُلِّلت البيانات باستخدام الحزمة المسمّاة» اختصار مقبول، أما «حلّلت البرمجية النوعية المقابلات» فليس كذلك، لأنها لم تفعل. ولا تُبلغ بالبرمجية بدل التقنية: فتسمية حزمة لا تخبر القارئ شيئًا عن أي اختبار شغّلت، والتقنية هي ما يهم. ولا تلصق مخرَج البرمجية الخام في الأطروحة، لأن المخرَج مصمَّم للمحلل وفصلُ النتائج مكتوب للقارئ؛ استخرج الأرقام التي تحتاجها في جدول صمّمته أنت.
ونقطة أكبر تختم هذا. فكفاءة البرمجيات تستحق الاكتساب وهي ليست المهارة التي تُقيَّم. فالممتحن لا يسأل أي حزمة استخدمت؛ بل يسأل لماذا اخترت تلك التقنية، وهل صحّت افتراضاتها، وما يعنيه حجم الأثر، وكيف تمسّ النتيجةُ سؤالك. والأطروحة يمكن بناؤها بأدوات متواضعة وحكمٍ ممتاز، ولا يمكن إنقاذها بأدوات متطورة وحكمٍ ضعيف. اختر الأداة التي تتيح لك إنفاق انتباهك على التحليل لا على الأداة.