Most searching fails before a single query is typed, because the searcher started from a topic rather than from a structure. A topic produces a query that returns either eleven thousand results or four, and in both cases the searcher concludes that the database is unhelpful when the problem was upstream.
A planned search has three properties that an improvised one lacks. It is systematic, meaning it covers the question's components deliberately rather than following whatever the first result suggests. It is reproducible, meaning another researcher given your protocol would retrieve substantially the same set. And it is reportable, meaning you can write a paragraph in your methodology that tells a reader exactly what was searched, when, and with what limits, which is a requirement at doctoral level and an expectation everywhere else.
The cost of planning is roughly two hours. The cost of not planning is measured differently: weeks of reading papers that turn out to be off-target, a persistent unease about what might have been missed, and an examiner's question you cannot answer. The asymmetry is the entire argument.
This framework covers how to decompose a question into searchable concepts, how to harvest the vocabulary the field actually uses, how to set boundaries you can defend, how to build and document a protocol, and how to know when to stop. It is my own synthesis, written in my own words and grounded in recognized scholarship.
معظم عمليات البحث تخفق قبل كتابة استعلام واحد، لأن الباحث انطلق من موضوع لا من بنية. والموضوع يُنتج استعلامًا يعيد إما أحد عشر ألف نتيجة أو أربعًا، وفي الحالتين يستنتج الباحث أن قاعدة البيانات غير مفيدة والمشكلة كانت أعلى المجرى.
وللبحث المخطَّط ثلاث خصائص يفتقدها المرتجَل. فهو منهجي، بمعنى أنه يغطي مكونات السؤال عن قصد لا يتبع ما توحي به أول نتيجة. وهو قابل للتكرار، بمعنى أن باحثًا آخر يُعطى بروتوكولك سيسترجع المجموعة نفسها جوهريًا. وهو قابل للإبلاغ، بمعنى أنك تستطيع كتابة فقرة في منهجيتك تخبر القارئ بالضبط ماذا بُحث ومتى وبأي حدود، وهذا مطلب في مستوى الدكتوراه وتوقُّعٌ في كل مكان آخر.
وكلفة التخطيط ساعتان تقريبًا. أما كلفة عدم التخطيط فتُقاس بمقياس آخر: أسابيع من قراءة أوراق تتبيّن خارج الهدف، وقلقٌ مستمر مما قد يكون فاتك، وسؤالُ ممتحنٍ لا تستطيع إجابته. وهذا اللاتماثل هو الحجة كلها.
ويغطي هذا الإطار كيف يُفكَّك السؤال إلى مفاهيم قابلة للبحث، وكيف تُحصَد المفردات التي يستخدمها الحقل فعلًا، وكيف تُضبَط حدود قابلة للدفاع، وكيف يُبنى بروتوكول ويُوثَّق، وكيف تعرف متى تتوقف. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
A searchable question is a set of two to four concept blocks. The blocks are joined by AND; the terms inside each block are joined by OR. Everything else in search technique is detail on top of that one structure.
The framework above, borrowed from evidence synthesis in health and adapted freely for business, is a prompt rather than a rule. Most business questions use three of the five. A study of whether flexible scheduling reduces turnover in small firms has a population, an intervention, and an outcome, and its context is the qualifier that decides transferability. A purely descriptive study may have only a population and a context.
The critical discipline is using two or three blocks, not five. Every additional block joined by AND multiplies the restriction, and a five-block query typically returns nothing. The practical approach is to search on the two blocks that are most distinctive, then filter the results by the remaining criteria during screening rather than in the query. Searching is for retrieval; screening is for precision, and confusing the two is the most common self-inflicted wound in literature searching.
Choosing which blocks are distinctive is a judgement call worth making explicitly. A block is distinctive when its terms appear in the title or abstract of the papers you want and rarely in others. Population terms are usually good blocks; outcome terms are usually good blocks; context terms such as a country name are often poor blocks, because many relevant papers mention the country only in the methods section where the database may not index it.
Once the blocks are fixed, write them down in the protocol as a table with one column per block. That table is the skeleton of every search string you will build, and it will be what you paste into the methodology chapter when the time comes.
السؤال القابل للبحث مجموعةٌ من كتلتين إلى أربع كتل مفهومية. تُوصَل الكتل بـAND؛ وتُوصَل المصطلحات داخل كل كتلة بـOR. وكل ما عدا ذلك في تقنية البحث تفصيلٌ فوق تلك البنية الواحدة.
والإطار أعلاه، المستعار من تركيب الأدلة في الصحة والمكيَّف بحرية للأعمال، تنبيهٌ لا قاعدة. فمعظم أسئلة الأعمال تستخدم ثلاثة من الخمسة. فدراسة هل تخفّض الجدولة المرنة دوران الموظفين في المنشآت الصغيرة لها مجتمع وتدخّل ومخرَج، وسياقها هو المقيّد الذي يقرر قابلية النقل. والدراسة الوصفية الخالصة قد يكون لها مجتمع وسياق فقط.
والانضباط الحاسم هو استخدام كتلتين أو ثلاث لا خمس. فكل كتلة إضافية موصولة بـAND تضاعف التقييد، والاستعلام ذو الخمس كتل يعيد لا شيء عادةً. والمقاربة العملية هي البحث بالكتلتين الأكثر تمييزًا، ثم ترشيح النتائج بالمعايير الباقية أثناء الفرز لا في الاستعلام. فالبحث للاسترجاع؛ والفرز للدقة، والخلط بينهما أشيع جرحٍ ذاتي في البحث في الأدبيات.
واختيار أي الكتل مميّزة حكمٌ يستحق أن يُتَّخذ صراحةً. فالكتلة مميّزة حين تظهر مصطلحاتها في عنوان أو ملخص الأوراق التي تريدها ونادرًا في غيرها. ومصطلحات المجتمع كتلٌ جيدة عادةً؛ ومصطلحات المخرَج كتلٌ جيدة عادةً؛ ومصطلحات السياق كاسم بلدٍ كتلٌ ضعيفة غالبًا، لأن أوراقًا كثيرة ذات صلة تذكر البلد في قسم المناهج فقط حيث قد لا تفهرسه القاعدة.
وبمجرد تثبيت الكتل، اكتبها في البروتوكول جدولًا بعمود لكل كتلة. فذلك الجدول هيكل كل سلسلة بحث ستبنيها، وهو ما ستلصقه في فصل المنهجية حين يحين الوقت.
A search finds the words you typed, not the ideas you meant. The gap between those two is where most missed literature lives, and closing it is a harvesting task rather than a thinking task.
Start from three known-relevant papers, found by any means including a supervisor's recommendation or a plain web search. These are the seed papers, and their function is not to be cited but to teach you the vocabulary. From each one, extract the title words, the author keywords, the database subject terms if shown, and the terms used repeatedly in the abstract. Ten minutes on three papers typically yields fifteen to twenty candidate terms, most of which you would not have invented.
Then expand each term along four axes. Synonyms: turnover, attrition, quit rate, separation, retention as its inverse. Spelling variants: organisation and organization, behaviour and behavior, labour and labor, which matter because databases do not always reconcile them. Word forms: manage, manages, managing, management, manager, which truncation handles in one stroke. Broader and narrower terms: flexible work as the broader, compressed week and remote work and flexitime as the narrower, since a search on the broad term alone will miss papers that only ever name the specific practice.
Watch for terms that changed over time, which is a systematic source of missed literature in fast-moving areas. Work published a decade ago may call the same practice telecommuting where current work calls it remote work or hybrid work. If your review covers a period during which the vocabulary shifted, both generations of terms belong in the block, and noticing the shift is itself a finding worth a sentence in the chapter.
Where a database offers a controlled vocabulary, meaning an official thesaurus of subject terms assigned by indexers, use it alongside your free-text terms rather than instead of them. Controlled terms catch papers whose authors used unusual language; free-text terms catch recent papers not yet indexed. The combination is more complete than either, and reporting that you used both is a mark of a careful search.
البحث يجد الكلمات التي كتبتها لا الأفكار التي قصدتها. والفجوة بين الاثنتين هي حيث تسكن معظم الأدبيات الفائتة، وإغلاقها مهمة حصادٍ لا مهمة تفكير.
ابدأ من ثلاث أوراق معلومة الصلة، تجدها بأي وسيلة بما فيها توصية مشرف أو بحث وب عادي. فهذه أوراق البذرة، ووظيفتها ليست أن يُستشهَد بها بل أن تعلّمك المفردات. ومن كل واحدة، استخرج كلمات العنوان، وكلمات المؤلف المفتاحية، ومصطلحات الموضوع في القاعدة إن عُرضت، والمصطلحات المستخدمة بتكرار في الملخص. وعشر دقائق على ثلاث أوراق تُنتج عادةً خمسة عشر إلى عشرين مصطلحًا مرشحًا، معظمها ما كنت لتبتكره.
ثم وسّع كل مصطلح على أربعة محاور. المترادفات: الدوران، والاستنزاف، ومعدل الترك، والانفصال، والبقاء بوصفه نقيضه. وتنويعات الهجاء: organisation وorganization، وbehaviour وbehavior، وlabour وlabor، وهي تهم لأن القواعد لا توفّق بينها دائمًا. وصيغ الكلمة: manage وmanages وmanaging وmanagement وmanager، وهذا يعالجه البتر بضربة واحدة. والمصطلحات الأعم والأخص: «العمل المرن» أعمّ، و«الأسبوع المضغوط» و«العمل عن بُعد» و«الوقت المرن» أخص، لأن البحث بالمصطلح الأعم وحده سيُفوّت أوراقًا لا تسمّي إلا الممارسة المحددة.
وانتبه لـالمصطلحات التي تغيّرت عبر الزمن، وهي مصدر منهجي للأدبيات الفائتة في المجالات سريعة الحركة. فعملٌ نُشر قبل عقد قد يسمّي الممارسة نفسها «العمل عن بُعد الاتصالي» بينما يسمّيها العمل الحالي «العمل عن بُعد» أو «العمل الهجين». وإن غطّت مراجعتك حقبةً تحوّلت فيها المفردات، فجيلا المصطلحات كلاهما مكانه الكتلة، وملاحظة التحول نفسها نتيجةٌ تستحق جملة في الفصل.
وحيث تعرض قاعدة مفردات مضبوطة، أي مكنزًا رسميًا لمصطلحات الموضوع يسنده المفهرسون، استخدمها إلى جانب مصطلحاتك الحرة لا بدلًا عنها. فالمصطلحات المضبوطة تلتقط أوراقًا استخدم مؤلفوها لغةً غير معتادة؛ والمصطلحات الحرة تلتقط أوراقًا حديثة لم تُفهرَس بعد. والجمع بينهما أكمل من أي منهما، والإبلاغ بأنك استخدمت الاثنتين علامةُ بحثٍ متأنٍّ.
Every review has boundaries. The difference between a good review and a thin one is whether the boundaries were chosen and stated or simply happened.
| Boundary | A defensible reason | A reason that fails |
|---|---|---|
| Years covered | The practice did not exist before this date | Older papers were harder to obtain |
| Language | The researcher can read only these languages, stated openly | Silence about language limits |
| Source types | Peer review required because the claims are causal | Grey literature excluded without saying so |
| Geography | The institutional setting changes the mechanism | Only local studies were convenient |
| Discipline | Adjacent fields define the construct differently | Adjacent fields were never searched |
| Sector | The outcome is sector-specific by definition | The first search happened to be in one sector |
The column on the right is not a list of sins. Every one of those reasons is real and most reviews are shaped by at least one of them. What separates an acceptable review from an unacceptable one is disclosure: a language restriction stated in the protocol is a limitation, and the same restriction unstated is a misrepresentation of coverage.
The date boundary deserves particular care because it is the one most often set thoughtlessly. The default of the last ten years is a convention with no justification behind it. A defensible start date is tied to an event: the publication of the foundational paper that introduced the construct, a regulatory change that altered the phenomenon, or a technology that did not previously exist. If you cannot name such an event, either the boundary is arbitrary and should be widened, or the arbitrariness should be stated.
The end date is a boundary too, and it is the one examiners notice. Record the date on which each search was last run, and rerun the searches shortly before submission. A review whose latest source is two years old invites the question of what appeared since, and the only good answer is a rerun that found nothing important.
Finally, distinguish search boundaries from screening criteria. A boundary is applied in the query or the database filter and shapes what is retrieved. A criterion is applied by reading and shapes what is kept. Keeping them separate in your protocol makes the eventual flow diagram straightforward and prevents the common error of describing screening decisions as though they were search limits.
لكل مراجعة حدود. والفرق بين مراجعة جيدة وأخرى ضحلة هو هل اختيرت الحدود وذُكرت أم أنها وقعت وحسب.
| الحد | سبب قابل للدفاع | سبب يخفق |
|---|---|---|
| السنوات المغطّاة | الممارسة لم تكن موجودة قبل هذا التاريخ | الأوراق الأقدم كان الحصول عليها أصعب |
| اللغة | الباحث لا يقرأ إلا هذه اللغات، مذكورًا صراحةً | الصمت عن حدود اللغة |
| أنواع المصادر | التحكيم مطلوب لأن الادّعاءات سببية | استبعاد الأدبيات الرمادية دون قول ذلك |
| الجغرافيا | الإطار المؤسسي يغيّر الآلية | الدراسات المحلية وحدها كانت ميسورة |
| التخصص | الحقول المجاورة تعرّف البناء تعريفًا مختلفًا | الحقول المجاورة لم تُبحث قط |
| القطاع | المخرَج خاص بالقطاع بالتعريف | البحث الأول صادف أن كان في قطاع واحد |
والعمود الأيمن ليس قائمة آثام. فكل واحد من تلك الأسباب حقيقي ومعظم المراجعات يشكّلها واحد منها على الأقل. وما يفصل مراجعةً مقبولة عن غير مقبولة هو الإفصاح: فتقييد اللغة المذكور في البروتوكول حدٌّ، والتقييد نفسه غير المذكور تحريفٌ للتغطية.
وحد التاريخ يستحق عناية خاصة لأنه الأكثر ضبطًا دون تفكير. فعُرف «السنوات العشر الأخيرة» عُرفٌ بلا تبرير خلفه. وتاريخ البدء القابل للدفاع مرتبط بحدث: نشر الورقة التأسيسية التي قدّمت البناء، أو تغيّر تنظيمي بدّل الظاهرة، أو تقنية لم تكن موجودة قبلًا. فإن لم تستطع تسمية حدث كهذا، فإما أن الحدّ اعتباطي وينبغي توسيعه، أو أن الاعتباطية ينبغي أن تُذكَر.
وتاريخ الانتهاء حدٌّ أيضًا، وهو الذي يلاحظه الممتحنون. سجّل التاريخ الذي شُغِّل فيه كل بحث آخر مرة، وأعِد تشغيل عمليات البحث قبيل التسليم. فالمراجعة التي أحدث مصادرها عمره سنتان تستدعي سؤالًا عمّا ظهر بعدها، والإجابة الجيدة الوحيدة إعادةُ تشغيلٍ لم تجد شيئًا مهمًا.
وأخيرًا، ميّز حدود البحث عن معايير الفرز. فالحدّ يُطبَّق في الاستعلام أو في مرشّح القاعدة ويشكّل ما يُسترجَع. والمعيار يُطبَّق بالقراءة ويشكّل ما يُبقى عليه. وفصلهما في بروتوكولك يجعل مخطط التدفق النهائي مباشرًا ويمنع الخطأ الشائع في وصف قرارات الفرز وكأنها حدود بحث.
The protocol is a one-page document written before the first search and updated as the search runs. It is the difference between a search you can report and a search you can only remember.
The protocol has seven entries and each takes a line or two. The question in its final wording. The concept blocks and the harvested terms in each. The databases to be searched, with a sentence on why those and not others. The boundaries, each with its reason. The inclusion and exclusion criteria to be applied at screening. The log, meaning a table with one row per search recording the database, the exact string, the date, and the number of results. And the deviations, meaning anything you changed after starting and why.
The log is the entry students skip and later regret. Record the exact string, not a description of it, because a string you paraphrase from memory is not the string you ran. Record the count, because the counts are what populate the flow diagram, and reconstructing them afterwards is impossible. Record the date, because databases are updated continuously and a rerun months later will return a different number, which is expected and must be explainable.
Testing recall is the step that turns a plausible search into a validated one, and it is quick. Take five papers you already know are relevant and check whether your search string retrieves all five. If it misses one, examine why: usually the paper uses a term absent from your blocks, and adding it fixes a systematic hole rather than a single omission. A string that retrieves all five known papers is not proof of completeness, but a string that misses two of five is proof of incompleteness.
Knowing when to stop rests on two signals used together. The first is saturation: new searches return papers you have already screened, and no new concepts appear in the ones that are new. The second is closure of the citation network: when you follow the reference lists of your key papers backwards and the papers citing them forwards, you arrive at works already in your set. When both signals hold, further searching has a low expected yield, and the correct response is to stop and write, with a note in the protocol that the search closed on that date.
البروتوكول وثيقة من صفحة تُكتَب قبل أول بحث وتُحدَّث أثناء تشغيله. وهو الفرق بين بحثٍ تستطيع الإبلاغ عنه وبحثٍ لا تستطيع إلا تذكّره.
وللبروتوكول سبعة مدخلات يأخذ كل منها سطرًا أو سطرين. السؤال بصياغته النهائية. وكتل المفاهيم والمصطلحات المحصودة في كل منها. والقواعد المزمع بحثها، مع جملة عن سبب هذه دون غيرها. والحدود، كلٌّ بسببه. ومعايير الإدراج والاستبعاد المزمع تطبيقها عند الفرز. والسجل، أي جدول بصف لكل بحث يسجّل القاعدة والسلسلة بالضبط والتاريخ وعدد النتائج. والانحرافات، أي كل ما غيّرته بعد البدء ولماذا.
والسجل هو المدخل الذي يتخطاه الطلاب ويندمون لاحقًا. سجّل السلسلة بالضبط لا وصفًا لها، لأن سلسلةً تعيد صياغتها من الذاكرة ليست السلسلة التي شغّلتها. وسجّل العدد، لأن الأعداد هي ما يملأ مخطط التدفق، وإعادة بنائها لاحقًا مستحيلة. وسجّل التاريخ، لأن القواعد تُحدَّث باستمرار وإعادة التشغيل بعد أشهر ستعيد رقمًا مختلفًا، وهذا متوقَّع ويجب أن يكون قابلًا للتفسير.
واختبار الاسترجاع هو الخطوة التي تحوّل بحثًا معقولًا إلى بحث مُتحقَّق منه، وهي سريعة. خذ خمس أوراق تعلم سلفًا أنها ذات صلة وافحص هل تسترجع سلسلتك الخمس كلها. فإن فاتت واحدة فافحص لماذا: عادةً تستخدم الورقة مصطلحًا غائبًا عن كتلك، وإضافته تسدّ ثغرة منهجية لا إغفالًا واحدًا. والسلسلة التي تسترجع الأوراق الخمس المعلومة ليست برهانًا على الاكتمال، لكن السلسلة التي تفوتها اثنتان من خمس برهانٌ على عدم الاكتمال.
ومعرفة متى تتوقف تقوم على إشارتين تُستخدَمان معًا. الأولى الإشباع: عمليات البحث الجديدة تعيد أوراقًا فرزتها سلفًا، ولا تظهر مفاهيم جديدة في الجديدة منها. والثانية انغلاق شبكة الاستشهاد: حين تتبع قوائم مراجع أوراقك المفتاحية إلى الخلف والأوراق المستشهدة بها إلى الأمام، تصل إلى أعمال في مجموعتك سلفًا. وحين تتحقق الإشارتان يكون للبحث الإضافي عائد متوقَّع منخفض، والاستجابة الصحيحة أن تتوقف وتكتب، مع ملاحظة في البروتوكول بأن البحث أُغلق في ذلك التاريخ.