Bibliometric analysis treats the publication record as data. Instead of reading papers to learn what they say, it counts and relates them to learn how a field is structured: who publishes, what is cited, which topics cluster, and how all of that has changed.
It answers a different question from a literature review, and the distinction matters. A review asks what is known about a phenomenon. A bibliometric analysis asks what the field looks like: its size, growth, concentration, intellectual groupings, and boundaries. Both are legitimate; using one where the other is needed is not.
Its advantage is scale. A researcher can read two hundred papers carefully; a bibliometric analysis can describe the structure of eight thousand, and it can identify the clusters within them from the data rather than from the researcher's prior categories. That last property is a real defence against the criticism that a review reproduced its author's assumptions.
Its limitation is equally clear and must be stated: it measures citation behaviour, not content. Papers cited together are related in some way, and the software cannot tell you in what way. Every bibliometric finding requires a reader to go and look at the papers before it means anything, and analyses that skip that step produce diagrams rather than knowledge.
This framework covers the two families of technique, building a defensible corpus, the standard performance indicators and what each does and does not measure, and the limits that govern interpretation. It is my own synthesis, written in my own words and grounded in recognized scholarship.
التحليل الببليومتري يعامل سجل النشر بوصفه بيانات. فبدل قراءة الأوراق لمعرفة ما تقول، يَعُدّها ويربطها لمعرفة كيف يتبنّى حقل: من ينشر، وما يُستشهَد به، وأي موضوعات تتعنقد، وكيف تغيّر ذلك كله.
وهو يجيب سؤالًا مختلفًا عن مراجعة الأدبيات، والتمييز يهم. فالمراجعة تسأل ما المعروف عن ظاهرة. والتحليل الببليومتري يسأل كيف يبدو الحقل: حجمه ونموه وتركّزه وتجمعاته الفكرية وحدوده. وكلاهما مشروع؛ واستخدام أحدهما حيث يلزم الآخر ليس كذلك.
وميزته الحجم. فالباحث يستطيع قراءة مئتَي ورقة بعناية؛ والتحليل الببليومتري يستطيع وصف بنية ثمانية آلاف، ويستطيع تحديد العناقيد داخلها من البيانات لا من فئات الباحث المسبقة. وتلك الخاصية الأخيرة دفاعٌ حقيقي ضد نقدٍ مفاده أن المراجعة أعادت إنتاج افتراضات مؤلفها.
وحدّه واضح بالقدر نفسه ويجب ذكره: إنه يقيس سلوك الاستشهاد لا المحتوى. فالأوراق المستشهَد بها معًا مرتبطة بطريقة ما، والبرمجية لا تستطيع إخبارك بأي طريقة. وكل نتيجة ببليومترية تتطلب أن يذهب قارئ وينظر في الأوراق قبل أن تعني شيئًا، والتحليلات التي تتخطى تلك الخطوة تُنتج مخططات لا معرفة.
ويغطي هذا الإطار عائلتَي التقنيات، وبناء مدونة قابلة للدفاع، ومؤشرات الأداء المعيارية وما يقيسه كلٌّ وما لا يقيسه، والحدود التي تحكم التفسير. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
Everything in bibliometrics is either counting or relating. Knowing which family a technique belongs to tells you what its output can support.
Performance analysis counts. Its outputs are the number of publications per year, the most productive authors, institutions, and countries, the most cited papers, the journals publishing most in the area, and the various indices combining productivity with impact. These describe activity, and they are descriptive statistics of a literature rather than findings about a phenomenon.
Their value in a thesis is contextual and real: establishing that a field is growing, that it is concentrated in a few outlets, that it is dominated by work from particular regions, or that a topic has emerged recently. All of these are things a reader would otherwise have to take on trust, and each takes one sentence supported by a count.
Science mapping relates. It uses the connections between records, which are citations and shared terms, to identify structure. Its outputs are clusters of works or authors that belong together, maps showing how those clusters relate, and views of how the structure has changed over time. This family is treated in detail in the next framework in this track.
Four mapping techniques recur and each answers a different question. Co-citation groups works that are cited together, revealing the intellectual foundations a field shares. Bibliographic coupling groups works that cite the same sources, revealing current research fronts. Co-word analysis groups terms that appear together, revealing thematic structure. And co-authorship maps collaboration between researchers, institutions, or countries.
The practical sequence for a thesis is to run performance analysis first, which is quick and gives you the field's shape, and then one or two mapping analyses chosen because they answer a question you actually have. Running all four because the software offers them produces four diagrams and no argument.
كل ما في الببليومتريا إما عدٌّ أو ربط. ومعرفةُ أي عائلة تنتمي إليها تقنية تخبرك بما يستطيع مخرَجها إسناده.
وتحليل الأداء يَعُدّ. ومخرجاته عدد المنشورات في السنة، وأكثر المؤلفين والمؤسسات والبلدان إنتاجًا، وأكثر الأوراق استشهادًا، والمجلات الأكثر نشرًا في المجال، والمؤشرات المتنوعة التي تجمع الإنتاجية بالأثر. وهذه تصف نشاطًا، وهي إحصاءات وصفية لأدبيات لا نتائج عن ظاهرة.
وقيمتها في أطروحة سياقية وحقيقية: إثباتُ أن حقلًا ينمو، أو أنه متركّز في منافذ قليلة، أو أن أعمالًا من مناطق بعينها تهيمن عليه، أو أن موضوعًا ظهر حديثًا. وكل هذه أشياء كان القارئ سيضطر إلى قبولها ثقةً لولا ذلك، وكلٌّ يستغرق جملةً يسندها عدد.
ورسم خرائط العلم يربط. فهو يستخدم الصلات بين السجلات، وهي الاستشهادات والمصطلحات المشتركة، لتحديد البنية. ومخرجاته عناقيدُ أعمال أو مؤلفين ينتمون معًا، وخرائط تُظهر كيف ترتبط تلك العناقيد، وعروضٌ لكيف تغيّرت البنية عبر الزمن. وتُعالَج هذه العائلة بالتفصيل في الإطار التالي من هذا المسار.
وأربع تقنيات رسم تتكرر وكلٌّ تجيب سؤالًا مختلفًا. الاستشهاد المشترك يجمّع الأعمال المستشهَد بها معًا، كاشفًا الأسس الفكرية التي يتشاركها حقل. والاقتران الببليوغرافي يجمّع الأعمال التي تستشهد بالمصادر نفسها، كاشفًا جبهات البحث الحالية. وتحليل الكلمات المشتركة يجمّع المصطلحات التي تظهر معًا، كاشفًا البنية الموضوعية. والتأليف المشترك يرسم التعاون بين الباحثين أو المؤسسات أو البلدان.
والتسلسل العملي لأطروحة تشغيلُ تحليل الأداء أولًا، وهو سريع ويعطيك شكل الحقل، ثم تحليل رسم أو تحليلين مختارين لأنهما يجيبان سؤالًا تملكه فعلًا. وتشغيلُ الأربعة لأن البرمجية تعرضها يُنتج أربعة مخططات ولا حجة.
Everything in a bibliometric analysis follows from the corpus. A badly defined corpus produces confident results about the wrong literature.
Define the corpus exactly as you would a systematic search: the database, the search string, the fields searched, the years, the document types, and the date of extraction. All six must be reported, because a bibliometric analysis is entirely determined by them and cannot be evaluated without them.
Two choices matter more here than in an ordinary search. Use one database, not several. Bibliometric analysis needs the citation relationships to be internally consistent, and merging records from two indexes creates duplicate and mismatched citation links that corrupt the network analyses. Choose the index with the better coverage of your field, say so, and note the coverage limitation.
Export the full records, including cited references and author keywords, not just the bibliographic details. Most of the mapping techniques need the reference lists, and an export without them will run only the counting analyses.
Cleaning is the step that determines whether the results mean anything, and it is the step most often skipped. Three problems recur. Author name variants: the same person appearing as J. Smith, John Smith, and Smith J, which splits their output across three nodes. Keyword synonyms: SME, SMEs, and small and medium enterprises counted as three distinct terms. And institution name variants, which are worse than author names because organizations rename and merge.
The tools support cleaning through a thesaurus file mapping variants to a canonical form. Building one takes a few hours for a corpus of a few thousand records, and it is not optional: an uncleaned analysis reports that the most common keyword is a term whose variants are scattered across six nodes, which is a finding about your data preparation rather than about the field.
Finally, report the corpus in the methodology with the same detail as a systematic search, including the numbers before and after each filter. A bibliometric chapter without a reproducible corpus definition cannot be checked, and it is the first thing a reader who knows the method will look for.
كل ما في التحليل الببليومتري يتبع من المدونة. والمدونةُ سيئة التعريف تُنتج نتائج واثقة عن الأدبيات الخطأ.
عرّف المدونة بالضبط كما تعرّف بحثًا منهجيًا: القاعدة، وسلسلة البحث، والحقول المبحوثة، والسنوات، وأنواع الوثائق، وتاريخ الاستخراج. والستة كلها يجب الإبلاغ بها، لأن التحليل الببليومتري تحدده كليًا ولا يمكن تقييمه بدونها.
وخياران يهمان هنا أكثر مما يهمان في بحث عادي. استخدم قاعدة واحدة لا عدة. فالتحليل الببليومتري يحتاج أن تكون علاقات الاستشهاد متسقة داخليًا، ودمجُ سجلات من فهرسين يخلق روابط استشهاد مكررة وغير متطابقة تُفسد تحليلات الشبكة. اختر الفهرس الأفضل تغطيةً لحقلك، وقُل ذلك، ودوّن حدّ التغطية.
صدّر السجلات كاملةً، بما فيها المراجع المستشهَد بها والكلمات المفتاحية للمؤلفين، لا البيانات الببليوغرافية فقط. فمعظم تقنيات الرسم تحتاج قوائم المراجع، والتصديرُ بدونها سيشغّل تحليلات العدّ فقط.
والتنظيف هو الخطوة التي تحدد هل تعني النتائج شيئًا، وهي الخطوة الأكثر تخطّيًا. وثلاث مشكلات تتكرر. تنويعات أسماء المؤلفين: الشخص نفسه يظهر بثلاث صيغ، وهذا يقسّم إنتاجه على ثلاث عقد. ومترادفات الكلمات المفتاحية: «منشأة صغيرة ومتوسطة» ومختصرها ومفردها تُعَدّ ثلاثة مصطلحات متمايزة. وتنويعات أسماء المؤسسات، وهي أسوأ من أسماء المؤلفين لأن المنظمات تُعاد تسميتها وتندمج.
والأدوات تسند التنظيف عبر ملف مكنز يربط التنويعات بصيغة معيارية. وبناؤه يستغرق ساعات قليلة لمدونة من بضعة آلاف سجل، وهو ليس اختياريًا: فتحليلٌ غير منظَّف يُبلغ بأن أشيع كلمة مفتاحية مصطلحٌ تتناثر تنويعاته على ست عقد، وهذه نتيجةٌ عن إعدادك للبيانات لا عن الحقل.
وأخيرًا، أبلغ بالمدونة في المنهجية بالتفصيل نفسه كبحث منهجي، بما فيه الأعداد قبل كل مرشِّح وبعده. فالفصل الببليومتري بلا تعريف مدونة قابل للتكرار لا يمكن فحصه، وهو أول ما سيبحث عنه قارئ يعرف المنهج.
Performance indicators are simple to compute and easy to misread. Each measures something specific, and none measures quality.
| Indicator | What it measures | What it does not |
|---|---|---|
| Publication count | Volume of output | Whether any of it mattered |
| Citation count | How often a work was referred to | Whether it was cited approvingly |
| Citations per publication | Average reach of an author's work | Distribution; one paper can carry the mean |
| The h-index | Sustained output with sustained citation | Career stage differences; it only rises |
| Journal impact measures | The average citation rate of a journal | The impact of any individual paper in it |
| Growth rate | How fast publication volume is changing | Whether the field is advancing |
Four cautions govern their use. Citations accumulate over time, so a paper from 2010 has had a decade to gather them and one from 2023 has not. Any comparison across years must normalise for this, and a top-cited list is always a list of older papers unless it is age-adjusted.
Citation practices differ by field, dramatically. A highly cited paper in one discipline would be unremarkable in another, so counts are comparable within a field and meaningless across fields.
Citation is not endorsement. Papers are cited to be criticised, to be corrected, and as examples of what not to do. A high count means a work was engaged with, which is a real signal, and it is not a quality measure.
Journal-level measures do not transfer to papers. The distribution of citations within any journal is highly skewed, so a paper in a high-impact journal may have no citations at all, and using the journal's number as a proxy for the paper is a well-documented error.
For a thesis, the honest and useful applications are contextual. Establishing the size and growth of a literature justifies a claim that a topic is emerging or established. Identifying the most cited works gives you a defensible list of papers to read closely, which is the most practical use of the whole method. Identifying the main outlets tells you where to search and where to publish. And identifying geographic or sectoral concentration can itself be the gap: a field studied almost entirely in one region is a field with an untested boundary.
مؤشرات الأداء بسيطة الحساب وسهلة الإساءة في القراءة. وكلٌّ يقيس شيئًا محددًا، ولا واحد يقيس الجودة.
| المؤشر | ما يقيسه | ما لا يقيسه |
|---|---|---|
| عدد المنشورات | حجم الإنتاج | هل كان أيٌّ منه مهمًا |
| عدد الاستشهادات | كم مرة أُشير إلى عمل | هل استُشهد به موافقةً |
| الاستشهادات لكل منشور | متوسط مدى عمل مؤلف | التوزيع؛ فورقة واحدة قد تحمل المتوسط |
| مؤشر هيرش | إنتاج مستدام باستشهاد مستدام | فروق مرحلة المسيرة؛ فهو يرتفع فقط |
| مقاييس أثر المجلات | متوسط معدل استشهاد مجلة | أثر أي ورقة مفردة فيها |
| معدل النمو | كم يتغير حجم النشر بسرعة | هل يتقدم الحقل |
وأربعة تحذيرات تحكم استخدامها. الاستشهادات تتراكم عبر الزمن، فورقةٌ من 2010 كان لديها عقدٌ لجمعها وورقةٌ من 2023 لم يكن لديها. وأي مقارنة عبر السنوات يجب أن تعيّر لهذا، وقائمةُ الأكثر استشهادًا دائمًا قائمةُ أوراق أقدم ما لم تُعدَّل بالعمر.
وممارسات الاستشهاد تختلف بالحقل، بدرجة كبيرة. فورقةٌ كثيرة الاستشهاد في تخصص ستكون عادية في آخر، فالأعداد قابلة للمقارنة داخل حقل وبلا معنى عبر الحقول.
والاستشهاد ليس تأييدًا. فالأوراق يُستشهَد بها للنقد وللتصحيح ومثالًا على ما لا يُفعَل. والعددُ العالي يعني أن عملًا جرى التعامل معه، وهذه إشارة حقيقية، وهو ليس مقياس جودة.
ومقاييس مستوى المجلة لا تنتقل إلى الأوراق. فتوزيع الاستشهادات داخل أي مجلة ملتوٍ بشدة، فقد لا يكون لورقة في مجلة عالية الأثر استشهادات البتة، واستخدامُ رقم المجلة وكيلًا عن الورقة خطأٌ موثَّق جيدًا.
وللأطروحة، التطبيقات الصادقة والنافعة سياقية. فإثباتُ حجم أدبيات ونموها يبرّر ادّعاءً بأن موضوعًا ناشئ أو راسخ. وتحديدُ أكثر الأعمال استشهادًا يعطيك قائمة أوراق قابلة للدفاع لقراءتها بإمعان، وهذا أعمل استخدام للمنهج كله. وتحديدُ المنافذ الرئيسة يخبرك أين تبحث وأين تنشر. وتحديدُ التركّز الجغرافي أو القطاعي قد يكون هو الفجوة نفسها: فحقلٌ يُدرَس كليًا تقريبًا في منطقة واحدة حقلٌ بحدّ غير مختبَر.
Bibliometric analysis produces confident-looking output from data that carries systematic biases. Six limits govern interpretation and all should be stated.
Database coverage. The analysis describes what the index contains, not the literature. Curated indexes under-represent regional journals, non-English work, books, and newer outlets, so a map built from one is a map of the indexed literature. State which index and acknowledge what it omits.
Language bias. English-language work is over-represented in every major index. For a study of a regionally specific phenomenon this can be severe, and the honest response is to say so and, where possible, to complement the analysis with a regional source.
Time lags. Recent papers have not accumulated citations, so any citation-based analysis systematically undervalues the newest work, which is often the work most relevant to an emerging topic. Bibliographic coupling, which uses reference lists rather than received citations, avoids this and is the better choice for identifying current fronts.
The Matthew effect. Highly cited papers attract further citations partly because they are visible, so citation counts partly measure prior visibility rather than merit. This is well documented and means top-cited lists are conservative: they reliably identify what a field considers foundational and are poor at identifying what is currently important.
Self-citation and citation circles inflate counts for some authors and groups. Most software can exclude self-citations and doing so is worth reporting.
And the central limit: citation is not content. Two papers cited together are related somehow, and the analysis cannot say how. A cluster is a hypothesis about intellectual structure, and it becomes a finding only when you read the central papers in it and can describe what they share.
For a thesis, three practices make bibliometric work defensible. Report the corpus fully, with the search string, database, filters, dates, and counts. Report the parameters: the software and version, the thresholds used, the clustering settings, and the cleaning performed, since different thresholds produce different maps from the same data. And read the papers: identify the most central works in each cluster and describe what they actually argue, because that reading is what converts a diagram into a contribution.
Used this way, bibliometrics is a genuine addition to a doctoral literature review: it establishes the field's shape objectively, it identifies the papers worth close reading from the data rather than from your prior expectations, and it can reveal a structural gap that reading alone would not surface. Used as a substitute for reading, it produces a chapter full of maps and empty of argument.
التحليل الببليومتري يُنتج مخرَجًا يبدو واثقًا من بيانات تحمل تحيّزات منهجية. وستة حدود تحكم التفسير وكلها ينبغي ذكرها.
تغطية القاعدة. فالتحليل يصف ما يحويه الفهرس، لا الأدبيات. والفهارس المنسّقة تقصّر في تمثيل المجلات الإقليمية والأعمال غير الإنجليزية والكتب والمنافذ الأحدث، فخريطةٌ مبنيّة من واحد خريطةٌ للأدبيات المفهرَسة. اذكر أي فهرس وأقرّ بما يُغفله.
تحيّز اللغة. فالأعمال بالإنجليزية ممثَّلة بإفراط في كل فهرس كبير. ولدراسة ظاهرة خاصة بإقليم قد يكون هذا شديدًا، والاستجابة الصادقة قولُ ذلك، وحيثما أمكن، تكميلُ التحليل بمصدر إقليمي.
الفوارق الزمنية. فالأوراق الحديثة لم تراكم استشهادات، فأي تحليل قائم على الاستشهاد يبخس منهجيًا أحدث الأعمال، وهي غالبًا الأعمال الأوثق صلةً بموضوع ناشئ. والاقتران الببليوغرافي، الذي يستخدم قوائم المراجع لا الاستشهادات المتلقاة، يتجنب هذا وهو الخيار الأفضل لتحديد الجبهات الحالية.
أثر ماثيو. فالأوراق كثيرة الاستشهاد تجتذب استشهادات أخرى جزئيًا لأنها مرئية، فأعدادُ الاستشهادات تقيس جزئيًا بروزًا سابقًا لا استحقاقًا. وهذا موثَّق جيدًا ويعني أن قوائم الأكثر استشهادًا محافظة: فهي تحدد بثبات ما يعدّه حقلٌ تأسيسيًا وهي ضعيفة في تحديد ما هو مهم حاليًا.
والاستشهاد الذاتي ودوائر الاستشهاد تضخّم الأعداد لبعض المؤلفين والمجموعات. ومعظم البرمجيات تستطيع استبعاد الاستشهادات الذاتية وفعلُ ذلك يستحق الإبلاغ.
والحد المركزي: الاستشهاد ليس محتوى. فورقتان مستشهَد بهما معًا مرتبطتان بطريقةٍ ما، ولا يستطيع التحليل قول كيف. والعنقود فرضيةٌ عن بنية فكرية، ويصير نتيجةً فقط حين تقرأ الأوراق المركزية فيه وتستطيع وصف ما تتشاركه.
وللأطروحة، ثلاث ممارسات تجعل العمل الببليومتري قابلًا للدفاع. أبلغ بالمدونة كاملةً، بسلسلة البحث والقاعدة والمرشِّحات والتواريخ والأعداد. وأبلغ بالمعاملات: البرمجية وإصدارها، والعتبات المستخدَمة، وإعدادات العنقدة، والتنظيف المُجرى، لأن عتبات مختلفة تُنتج خرائط مختلفة من البيانات نفسها. واقرأ الأوراق: حدّد أكثر الأعمال مركزيةً في كل عنقود وصِف ماذا تحاجّ به فعلًا، لأن تلك القراءة هي ما يحوّل مخططًا إلى إسهام.
ومستخدَمةً هكذا، تكون الببليومتريا إضافةً حقيقية لمراجعة أدبيات في الدكتوراه: فهي تثبّت شكل الحقل موضوعيًا، وتحدد الأوراق التي تستحق القراءة المتأنية من البيانات لا من توقعاتك المسبقة، وتستطيع كشف فجوة بنيوية لم تكن القراءة وحدها لتُظهرها. ومستخدَمةً بديلًا عن القراءة، تُنتج فصلًا مليئًا بالخرائط خاليًا من الحجة.