A systematic review answers one narrowly defined question by a procedure fixed in advance and reported in full. Its distinguishing feature is not thoroughness, which any good review has, but reproducibility: another researcher given the protocol should arrive at substantially the same included set and therefore at the same conclusion.
That property is what makes the output evidence about a body of literature rather than an account of one reader's engagement with it. A narrative review says here is what I read and what I think it means. A systematic review says here is everything that met these criteria, and here is what it collectively shows. The second is a stronger claim and requires a proportionally stronger procedure.
PRISMA, which stands for preferred reporting items for systematic reviews and meta-analyses, is the guideline that governs how such a review is reported. It is worth being precise about what it is: a reporting standard, not a method. It tells you what a reader needs to be told; it does not tell you how to do the work. A review can follow every reporting item and still be a poor review, and a good review that omits the reporting items is one a reader cannot evaluate.
This framework covers the six stages of a systematic review, what the reporting guideline requires and what it leaves open, how quality appraisal works and why it is separate from screening, how data extraction is designed, the realistic workload, and how the standard adapts to business research where it originated in health. It is my own synthesis, written in my own words and grounded in recognized scholarship.
المراجعة المنهجية تجيب سؤالًا واحدًا محدَّدًا بضيق بإجراء يُثبَّت سلفًا ويُبلَّغ عنه كاملًا. وسمتها المميّزة ليست الاستقصاء، فهذا تملكه كل مراجعة جيدة، بل القابلية للتكرار: فباحثٌ آخر يُعطى البروتوكول ينبغي أن يصل إلى المجموعة المدرجة نفسها جوهريًا وبالتالي إلى الخاتمة نفسها.
وتلك الخاصية هي ما يجعل المخرَج دليلًا عن جسم من الأدبيات لا عرضًا لتعامل قارئ واحد معه. فالمراجعة السردية تقول «هذا ما قرأت وهذا ما أظنه يعنيه». والمراجعة المنهجية تقول «هذا كل ما استوفى هذه المعايير، وهذا ما يُظهره مجتمعًا». والثاني ادّعاء أقوى ويتطلب إجراءً أقوى بالتناسب.
وبريزما، وهي اختصار «بنود الإبلاغ المفضّلة للمراجعات المنهجية والتحليل البعدي»، هي الدليل الذي يحكم كيف يُبلَّغ عن مراجعة كهذه. ويستحق الأمر دقةً في ما هي: معيار إبلاغ لا منهج. فهي تخبرك بما يحتاج القارئ أن يُخبَر به؛ ولا تخبرك كيف تؤدي العمل. ويمكن لمراجعة أن تتبع كل بند إبلاغ وتظل مراجعةً ضعيفة، والمراجعة الجيدة التي تُغفل بنود الإبلاغ مراجعةٌ لا يستطيع القارئ تقييمها.
ويغطي هذا الإطار مراحل المراجعة المنهجية الست، وما يشترطه دليل الإبلاغ وما يتركه مفتوحًا، وكيف يعمل تقييم الجودة ولماذا هو منفصل عن الفرز، وكيف يُصمَّم استخلاص البيانات، وحجم العمل الواقعي، وكيف يتكيّف المعيار مع بحوث الأعمال وقد نشأ في الصحة. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
A systematic review is a pipeline. Each stage takes a defined input, produces a defined output, and cannot be run out of order without damaging the ones after it.
Stage one, the protocol, fixes the question, the inclusion and exclusion criteria, the databases and search strings, the appraisal tool, the extraction fields, and the intended synthesis method. Writing all of this before searching is the entire point, because every one of these decisions could otherwise be influenced by the results. A protocol is three to six pages and takes a week of thinking.
Stage two, identification, runs the searches and records what each returned, then removes duplicates. Its outputs are numbers, and those numbers appear unchanged at the top of the flow diagram. Supplementary sources such as citation chasing and hand-searching are logged separately here rather than merged into the database counts.
Stage three, screening, applies the criteria in two passes and records an exclusion reason for every full-text rejection. Where two screeners are available they work independently and agreement is measured; where only one is available, the honest alternative is a documented consistency check on a sample.
Stage four, appraisal, assesses how much confidence each included study earns. This is separate from screening: screening asks does this study belong in the review, appraisal asks how much weight it deserves within it. Studies are rarely excluded on appraisal alone; more often the appraisal becomes a variable in the synthesis, so that a pattern found only in weak studies is reported as such.
Stage five, extraction, pulls the same fields from every included study into one table. Designing that table well is what makes stage six possible, and redesigning it after extracting forty studies is the most common avoidable rework in the whole process.
Stage six, synthesis, combines the extracted evidence, statistically if the studies permit pooling and narratively if they do not, and reports the result with its limitations. The synthesis answers the protocol's question, and if it answers a different question, something went wrong earlier.
المراجعة المنهجية خطُّ إنتاج. كل مرحلة تأخذ مدخلًا محدَّدًا، وتُنتج مخرَجًا محدَّدًا، ولا يمكن تشغيلها خارج الترتيب دون إضرار بما بعدها.
المرحلة الأولى، البروتوكول، تثبّت السؤال ومعايير الإدراج والاستبعاد والقواعد وسلاسل البحث وأداة التقييم وحقول الاستخلاص ومنهج التركيب المزمع. وكتابة هذا كله قبل البحث هي المغزى كله، لأن كل واحد من هذه القرارات قد يتأثر بالنتائج لولا ذلك. والبروتوكول ثلاث إلى ست صفحات ويستغرق أسبوعًا من التفكير.
والمرحلة الثانية، التعرّف، تشغّل عمليات البحث وتسجّل ما أعادته كل واحدة، ثم تزيل التكرار. ومخرجاتها أرقام، وتلك الأرقام تظهر دون تغيير في أعلى مخطط التدفق. والمصادر التكميلية كتتبّع الاستشهادات والبحث اليدوي تُسجَّل هنا منفصلةً لا مدموجةً في أعداد القواعد.
والمرحلة الثالثة، الفرز، تطبّق المعايير في مرورين وتسجّل سبب استبعاد لكل رفض في النص الكامل. وحيث يتوفر فارزان يعملان مستقلين ويُقاس التوافق؛ وحيث لا يتوفر إلا واحد، فالبديل الصادق فحصُ اتساق موثَّق على عيّنة.
والمرحلة الرابعة، التقييم، تقدّر كم ثقةً تستحق كل دراسة مدرجة. وهذه منفصلة عن الفرز: فالفرز يسأل هل تنتمي هذه الدراسة إلى المراجعة، والتقييم يسأل كم وزنًا تستحق داخلها. ونادرًا ما تُستبعَد دراسات على التقييم وحده؛ والأشيع أن يصير التقييم متغيرًا في التركيب، فيُبلَّغ عن نمط وُجد في الدراسات الضعيفة وحدها بوصفه كذلك.
والمرحلة الخامسة، الاستخلاص، تسحب الحقول نفسها من كل دراسة مدرجة إلى جدول واحد. وتصميم ذلك الجدول جيدًا هو ما يجعل المرحلة السادسة ممكنة، وإعادة تصميمه بعد استخلاص أربعين دراسة أشيعُ عمل معاد يمكن تجنبه في العملية كلها.
والمرحلة السادسة، التركيب، تجمع الأدلة المستخلَصة، إحصائيًا إن سمحت الدراسات بالتجميع وسرديًا إن لم تسمح، وتُبلغ عن النتيجة بحدودها. والتركيب يجيب سؤال البروتوكول، فإن أجاب سؤالًا آخر فشيءٌ ما أخفق مبكرًا.
The guideline is a checklist of items a report should contain and a flow diagram of how records moved through the review. Understanding its scope prevents both under-use and overclaim.
The checklist covers the whole report, section by section. In the title and abstract it asks that the review be identified as such and that the abstract carry a structured summary. In the introduction it asks for the rationale and the objectives stated as a specified question. In the methods it asks for the eligibility criteria, the information sources with dates, the full search strategy for at least one database, the selection process, the data collection process, the outcomes sought, the appraisal method, and the synthesis method.
In the results it asks for the flow of records, the characteristics of included studies, the appraisal results, the individual study results, and the synthesis. In the discussion it asks for an interpretation, the limitations of the evidence and of the review process, and the implications. It closes with items on registration, funding, competing interests, and availability of data and code.
The flow diagram is the element readers check first. It shows records identified by source, duplicates removed, records screened, records excluded at screening, reports sought for retrieval, reports not retrieved, reports assessed for eligibility, reports excluded with reasons, and studies included. The arithmetic must reconcile at every level, and it is the fastest available signal of whether a review was carefully conducted.
Three things the guideline does not do are worth stating. It does not specify how many databases to search or which ones. It does not require any particular appraisal tool. And it does not require dual screening, though the broader systematic review tradition assumes it. These are methodological choices you must make and justify, and the guideline's role is to make you disclose what you chose.
For a doctoral thesis the practical use is to treat the checklist as a completeness test for the methodology section rather than as a form to submit. Read down the items, and for each one either point to where in your chapter it is addressed or decide consciously that it does not apply. Items you cannot point to are gaps, and finding them before an examiner does costs an hour.
الدليل قائمةُ فحصٍ ببنود ينبغي أن يحويها التقرير ومخططُ تدفق لكيفية حركة السجلات عبر المراجعة. وفهم نطاقه يمنع قلة الاستعمال والادّعاء الزائد معًا.
وقائمة الفحص تغطي التقرير كله، قسمًا قسمًا. ففي العنوان والملخص تطلب أن تُعرَّف المراجعة كذلك وأن يحمل الملخص موجزًا مهيكَلًا. وفي المقدمة تطلب المسوِّغ والأهداف مذكورةً سؤالًا محدَّدًا. وفي المناهج تطلب معايير الأهلية، ومصادر المعلومات بتواريخها، واستراتيجية البحث كاملةً لقاعدة واحدة على الأقل، وعملية الاختيار، وعملية جمع البيانات، والمخرجات المطلوبة، ومنهج التقييم، ومنهج التركيب.
وفي النتائج تطلب تدفق السجلات، وخصائص الدراسات المدرجة، ونتائج التقييم، ونتائج الدراسات فرادى، والتركيب. وفي المناقشة تطلب تفسيرًا، وحدود الأدلة وحدود عملية المراجعة، والدلالات. وتُختَم ببنود عن التسجيل والتمويل وتضارب المصالح وإتاحة البيانات والشفرة.
ومخطط التدفق هو العنصر الذي يفحصه القرّاء أولًا. وهو يُظهر السجلات المعرَّفة بالمصدر، والتكرارات المزالة، والسجلات المفروزة، والمستبعَدة عند الفرز، والتقارير المطلوب استرجاعها، وغير المسترجَعة، والمقيَّمة للأهلية، والمستبعَدة بأسبابها، والدراسات المدرجة. والحساب يجب أن يتوافق في كل مستوى، وهو أسرع إشارة متاحة إلى هل أُجريت المراجعة بعناية.
وثلاثة أشياء لا يفعلها الدليل تستحق الذكر. فهو لا يحدد كم قاعدة تُبحَث ولا أيّها. ولا يشترط أداة تقييم بعينها. ولا يشترط فرزًا مزدوجًا، وإن كان تقليد المراجعة المنهجية الأوسع يفترضه. وهذه خيارات منهجية عليك اتخاذها وتبريرها، ودور الدليل أن يجعلك تُفصح عمّا اخترت.
وللأطروحة، الاستخدام العملي أن تعامل القائمة اختبارَ اكتمال لقسم المنهجية لا نموذجًا يُقدَّم. اقرأ البنود، ولكل واحد إما أشِر إلى موضع معالجته في فصلك أو قرّر واعيًا أنه لا ينطبق. والبنود التي لا تستطيع الإشارة إليها فجوات، وإيجادها قبل أن يجدها ممتحن يكلّف ساعة.
These two stages are where a systematic review does its real intellectual work, and where reviews most often become mechanical instead.
Appraisal asks how much confidence each study earns, on dimensions that vary by design. For a survey the questions concern the sampling frame and response rate, the validity evidence for the instrument, and whether the analysis matched the measurement level. For a qualitative study they concern the appropriateness of the sampling, the transparency of the analysis, and whether the interpretation is traceable to the data. For an experiment they concern allocation, control, and attrition.
Published appraisal tools exist for each design family and using one is better than inventing your own, because a named tool is a standard a reader recognises. What matters more than the choice of tool is that the appraisal is applied consistently and reported per study, usually as a table with one row per study and one column per criterion. That table lets a reader see for themselves whether the strongest finding rests on the strongest studies.
Resist the temptation to reduce appraisal to a score. Summing criteria into a single number implies that the criteria are commensurable and equally weighted, which they are not, and a study weak on one critical dimension is not compensated by strength on four minor ones. Report the profile rather than the total, and if you must categorise, use three broad bands with the reasoning stated.
Extraction is the design of a table. Its columns are decided in the protocol and typically include the citation, the country and sector, the design, the sample size and composition, the constructs and how they were measured, the analysis technique, the main finding on your question, the effect size or its qualitative equivalent, the appraisal result, and a free-text note. Extract the same fields from every study even when a study reports nothing for one, because an empty cell is itself a finding about the literature.
Two practical rules save considerable rework. Pilot the extraction form on five studies before running it on the set, because the columns you think you need and the columns you actually need differ in ways only contact with real papers reveals. And extract what the paper says, not what you conclude from it, keeping interpretation in the note column, so that the table remains a record of the evidence and the synthesis remains visibly your own work.
هاتان المرحلتان هما حيث تؤدي المراجعة المنهجية عملها الفكري الحقيقي، وحيث تصير المراجعات آليةً بدلًا من ذلك في أغلب الأحيان.
والتقييم يسأل كم ثقةً تستحق كل دراسة، على أبعاد تتباين بالتصميم. فللمسح تتعلق الأسئلة بإطار المعاينة ومعدل الاستجابة، ودليل صدق الأداة، وهل طابق التحليلُ مستوى القياس. وللدراسة النوعية تتعلق بملاءمة المعاينة، وشفافية التحليل، وهل التفسير قابل للتتبّع إلى البيانات. وللتجربة تتعلق بالتخصيص والضبط والتسرّب.
وثمة أدوات تقييم منشورة لكل عائلة تصميم واستخدام واحدة خيرٌ من ابتكار أداتك، لأن الأداة المسمّاة معيارٌ يعرفه القارئ. وما يهم أكثر من اختيار الأداة أن يُطبَّق التقييم باتساق ويُبلَّغ عنه لكل دراسة، جدولًا عادةً بصف لكل دراسة وعمود لكل معيار. فذلك الجدول يتيح للقارئ أن يرى بنفسه هل تستند أقوى نتيجة إلى أقوى الدراسات.
وقاوم إغراء اختزال التقييم إلى درجة. فجمع المعايير في رقم واحد يوحي بأنها متكافئة ومتساوية الوزن، وهي ليست كذلك، والدراسة الضعيفة في بُعد حرج واحد لا تعوّضها قوةٌ في أربعة أبعاد ثانوية. أبلغ عن الملف لا عن المجموع، وإن وجب التصنيف فاستخدم ثلاث فئات عريضة مع ذكر التعليل.
والاستخلاص تصميمُ جدول. وأعمدته تُقرَّر في البروتوكول وتشمل عادةً الاستشهاد، والبلد والقطاع، والتصميم، وحجم العيّنة وتركيبها، والبناءات وكيف قيست، وتقنية التحليل، والنتيجة الرئيسة على سؤالك، وحجم الأثر أو مكافئه النوعي، ونتيجة التقييم، وملاحظة نصية حرة. واستخلص الحقول نفسها من كل دراسة حتى حين لا تُبلغ دراسةٌ عن شيء لأحدها، لأن الخانة الفارغة نفسها نتيجةٌ عن الأدبيات.
وقاعدتان عمليتان توفّران عملًا معادًا كثيرًا. جرّب نموذج الاستخلاص على خمس دراسات قبل تشغيله على المجموعة، لأن الأعمدة التي تظن أنك تحتاجها والأعمدة التي تحتاجها فعلًا تختلفان بطرق لا يكشفها إلا التماس مع أوراق حقيقية. واستخلص ما تقوله الورقة لا ما تستنتجه أنت منها، مبقيًا التفسير في عمود الملاحظة، ليبقى الجدول سجلًا للأدلة ويبقى التركيب عملك أنت بوضوح.
A systematic review is a project, not a chapter written on the way to one. Knowing the real numbers before starting is what prevents the decision from being made by accident.
| Stage | Typical effort | What drives it |
|---|---|---|
| Protocol | One to two weeks | How settled the question is before you start |
| Searching and deduplication | One week | Number of databases and syntax translation |
| Title and abstract screening | One to three weeks | Roughly one to two hundred records an hour |
| Full-text screening | Two to four weeks | Obtaining files is often slower than reading them |
| Appraisal and extraction | Three to eight weeks | About one to two hours per included study |
| Synthesis and writing | Four to eight weeks | How heterogeneous the included set turned out to be |
The totals land between four and six months of concentrated work for a review of forty to eighty included studies, and longer if the searches return thousands of records or the field is methodologically diverse. That is a substantial commitment and it is why the decision belongs in the research plan rather than in the writing schedule.
Adapting the standard to business research requires three specific accommodations. First, heterogeneity is the norm: business studies of the same relationship use different constructs, instruments, and levels of analysis far more often than clinical studies do, so narrative synthesis rather than meta-analysis is the usual outcome and should be planned for rather than discovered.
Second, appraisal tools were built for health designs and transfer imperfectly. Survey and case study research needs criteria about construct validity, common method variance, sampling frames, and context specification that clinical tools do not contain. Using a management-appropriate tool, or adapting one and saying that you did, is better than forcing a mismatch.
Third, grey literature matters more. Consultancy reports, regulator publications, and industry statistics often hold the only evidence on a practice, and excluding them for lack of peer review can systematically bias a business review towards academically visible but practically marginal topics. Decide a policy, state it, and if you include grey sources, appraise them explicitly rather than silently.
The final judgement is proportionality. If the question genuinely requires a reproducible answer, and the field has enough comparable studies to answer it, and the timeline allows, a systematic review is the right instrument and produces an output that stands on its own. If any of those three conditions fails, a structured narrative review with a documented search delivers most of the transparency at a quarter of the cost, and describing it honestly is what makes it credible.
المراجعة المنهجية مشروعٌ لا فصلٌ يُكتَب في الطريق إلى مشروع. ومعرفة الأرقام الحقيقية قبل البدء هي ما يمنع اتخاذ القرار مصادفةً.
| المرحلة | الجهد النموذجي | ما يحرّكه |
|---|---|---|
| البروتوكول | أسبوع إلى أسبوعين | مدى استقرار السؤال قبل البدء |
| البحث وإزالة التكرار | أسبوع | عدد القواعد وترجمة الصياغة |
| فرز العنوان والملخص | أسبوع إلى ثلاثة | نحو مئة إلى مئتي سجل في الساعة |
| فرز النص الكامل | أسبوعان إلى أربعة | الحصول على الملفات أبطأ غالبًا من قراءتها |
| التقييم والاستخلاص | ثلاثة إلى ثمانية أسابيع | نحو ساعة إلى ساعتين لكل دراسة مدرجة |
| التركيب والكتابة | أربعة إلى ثمانية أسابيع | مدى تباين المجموعة المدرجة كما تبيّن |
والمجاميع تقع بين أربعة وستة أشهر من العمل المركَّز لمراجعة من أربعين إلى ثمانين دراسة مدرجة، وأطول إن أعادت عمليات البحث آلاف السجلات أو كان الحقل متنوعًا منهجيًا. وهذا التزام كبير ولهذا يخص القرارُ خطةَ البحث لا جدول الكتابة.
وتكييف المعيار مع بحوث الأعمال يتطلب ثلاثة تعديلات محددة. أولًا، التباين هو القاعدة: فدراسات الأعمال للعلاقة نفسها تستخدم بناءات وأدوات ومستويات تحليل مختلفة أكثر بكثير من الدراسات السريرية، فالتركيب السردي لا التحليل البعدي هو المخرَج المعتاد وينبغي التخطيط له لا اكتشافه.
وثانيًا، أدوات التقييم بُنيت لتصاميم صحية وتنتقل بنقصان. فبحوث المسح ودراسة الحالة تحتاج معايير عن صدق البناء وتباين المنهج المشترك وأطر المعاينة وتحديد السياق لا تحويها الأدوات السريرية. واستخدام أداة ملائمة للإدارة، أو تكييف واحدة وقول أنك فعلت، خيرٌ من فرض عدم تطابق.
وثالثًا، الأدبيات الرمادية تهم أكثر. فتقارير الاستشارات ومنشورات الجهات التنظيمية وإحصاءات القطاعات كثيرًا ما تحمل الدليل الوحيد على ممارسة، واستبعادها لانعدام التحكيم قد يحيّز مراجعةَ أعمالٍ منهجيًا نحو موضوعات مرئية أكاديميًا وهامشية عمليًا. قرّر سياسةً، واذكرها، وإن أدرجت مصادر رمادية فقيّمها صراحةً لا صمتًا.
والحكم الأخير التناسب. فإن كان السؤال يتطلب فعلًا إجابةً قابلة للتكرار، وكان للحقل دراسات قابلة للمقارنة تكفي لإجابته، وسمح الجدول الزمني، فالمراجعة المنهجية هي الأداة الصحيحة وتُنتج مخرَجًا قائمًا بذاته. وإن أخفق أي من الشروط الثلاثة، فالمراجعة السردية المهيكَلة ببحث موثَّق تقدّم معظم الشفافية بربع الكلفة، ووصفها بصدق هو ما يجعلها موثوقة.