Science mapping identifies the structure of a field by analysing the connections between its publications. Where performance analysis counts, mapping relates, and the output is a network in which groups of related work become visible.
Its value in a doctoral literature review is specific and worth stating clearly. The clusters are derived from the data rather than from your prior categories. A researcher organizing a literature by hand groups papers according to the distinctions they already believe matter; a mapping analysis groups them according to how the field itself cites and phrases things. When the two agree, you have evidence that your framework matches the field. When they disagree, the disagreement is informative and is often the most interesting thing in the analysis.
The technical core is a network: items are nodes, relations are edges, and an algorithm identifies groups of nodes more connected to each other than to the rest. That is the whole idea, and every technique in this framework is a different choice about what the nodes are and what counts as a relation.
The interpretive core is a discipline: a cluster is a hypothesis, not a finding. The software tells you that a group of papers is connected; it cannot tell you what they share. Only reading the central papers in a cluster turns it into a claim about the field, and an analysis that stops at the diagram has stopped one step short.
This framework covers the four relations and when each applies, the network concepts you need to read a map, how thresholds and parameters change the result, and how to interpret and report. It is my own synthesis, written in my own words and grounded in recognized scholarship.
رسم خرائط العلم يحدد بنية حقل بتحليل الصلات بين منشوراته. وحيث يَعُدّ تحليلُ الأداء، يربط الرسمُ، والمخرَج شبكةٌ تصير فيها مجموعات الأعمال المترابطة مرئية.
وقيمته في مراجعة أدبيات في الدكتوراه محددة وتستحق الذكر بوضوح. فالعناقيد مشتقة من البيانات لا من فئاتك المسبقة. فالباحث الذي ينظّم أدبيات يدويًا يجمّع الأوراق بحسب التمييزات التي يعتقد سلفًا أنها تهم؛ وتحليلُ الرسم يجمّعها بحسب كيف يستشهد الحقل نفسه ويصوغ الأشياء. وحين يتفق الاثنان، يكون لديك دليل أن إطارك يطابق الحقل. وحين يختلفان، يكون الاختلاف مفيدًا وهو غالبًا أطرف ما في التحليل.
والنواة التقنية شبكة: العناصر عقد، والعلاقات حواف، وخوارزمية تحدد مجموعات العقد الأشد اتصالًا ببعضها منها ببقية الشبكة. وتلك الفكرة كلها، وكل تقنية في هذا الإطار خيارٌ مختلف عمّا تكونه العقد وما يُعدّ علاقة.
والنواة التفسيرية انضباط: العنقود فرضيةٌ لا نتيجة. فالبرمجية تخبرك أن مجموعة أوراق متصلة؛ ولا تستطيع إخبارك بما تتشاركه. وقراءةُ الأوراق المركزية في عنقود وحدها تحوّله إلى ادّعاء عن الحقل، والتحليلُ الذي يتوقف عند المخطط قد توقف خطوةً قبل.
ويغطي هذا الإطار العلاقات الأربع ومتى تنطبق كلٌّ، ومفاهيم الشبكة التي تحتاجها لقراءة خريطة، وكيف تغيّر العتبات والمعاملات النتيجة، وكيف يُفسَّر ويُبلَّغ. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
Each technique defines a relation between records. Choosing among them is choosing what question you are asking about the field.
Co-citation connects two works when a third work cites both. It maps the intellectual foundations a field shares, because works consistently cited together are treated by the field as belonging together. Its properties follow from its definition. It looks backwards, since the cited works are older than the citing ones, and it therefore identifies the classics rather than the frontier. And it requires accumulated citations, so recent work cannot appear.
Bibliographic coupling connects two works when they cite the same source. It maps current research fronts, because works drawing on the same literature are working on related problems now. Its properties are the mirror image. It looks at the present and it works for very recent papers, because a paper's reference list exists the day it is published. For a doctoral student asking what is being worked on now, this is usually the more useful of the two, and it is the less commonly used.
Co-word analysis connects terms that appear together in titles, abstracts, or keyword lists. It maps thematic structure directly rather than through citation, which makes it the most interpretable of the four and also the most sensitive to cleaning, since unmerged synonyms fragment a theme across several nodes.
Co-authorship connects researchers, institutions, or countries who publish together. It maps collaboration rather than ideas, and in a thesis its main use is contextual: showing that a field is dominated by a few groups, or that a region participates little, which can itself be the basis of a gap claim.
Two further techniques appear in the literature and are worth knowing by name. Direct citation networks connect works where one cites the other, which is the most direct relation and requires a large corpus to be informative. And overlay visualisations colour a map by publication year, showing which parts of the structure are old and which are emerging, which is often the single most useful view for identifying where a field is moving.
كل تقنية تعرّف علاقةً بين السجلات. والاختيار بينها اختيارٌ لأي سؤال تسأله عن الحقل.
والاستشهاد المشترك يصل عملين حين يستشهد بهما عملٌ ثالث. وهو يرسم الأسس الفكرية التي يتشاركها حقل، لأن الأعمال المستشهَد بها معًا باتساق يعاملها الحقل بوصفها تنتمي معًا. وخصائصه تتبع من تعريفه. فهو ينظر خلفًا، لأن الأعمال المستشهَد بها أقدم من المستشهِدة، ولذلك يحدد الكلاسيكيات لا الجبهة. وهو يتطلب استشهادات متراكمة، فلا تستطيع الأعمال الحديثة الظهور.
والاقتران الببليوغرافي يصل عملين حين يستشهدان بالمصدر نفسه. وهو يرسم جبهات البحث الحالية، لأن الأعمال التي تستمد من الأدبيات نفسها تعمل على مشكلات مترابطة الآن. وخصائصه صورةٌ معكوسة. فهو ينظر إلى الحاضر ويعمل للأوراق الحديثة جدًا، لأن قائمة مراجع ورقة توجد يوم نشرها. ولطالب دكتوراه يسأل ما الذي يُعمَل عليه الآن، هذا عادةً الأنفع من الاثنين، وهو الأقل استعمالًا.
وتحليل الكلمات المشتركة يصل المصطلحات التي تظهر معًا في العناوين أو الملخصات أو قوائم الكلمات المفتاحية. وهو يرسم البنية الموضوعية مباشرةً لا عبر الاستشهاد، وهذا يجعله الأقبل للتفسير من الأربعة والأشد حساسيةً للتنظيف أيضًا، لأن المترادفات غير المدموجة تجزّئ موضوعًا على عدة عقد.
والتأليف المشترك يصل الباحثين أو المؤسسات أو البلدان الذين ينشرون معًا. وهو يرسم التعاون لا الأفكار، وفي أطروحة يكون استخدامه الرئيس سياقيًا: إظهارُ أن مجموعات قليلة تهيمن على حقل، أو أن منطقةً تشارك قليلًا، وهذا قد يكون بنفسه أساس ادّعاء فجوة.
وتقنيتان أخريان تظهران في الأدبيات وتستحقان المعرفة بالاسم. شبكات الاستشهاد المباشر تصل الأعمال حيث يستشهد أحدها بالآخر، وهي أكثر العلاقات مباشرةً وتتطلب مدونة كبيرة لتكون مفيدة. والتصويرات الطبقية تلوّن خريطةً بسنة النشر، مُظهِرةً أي أجزاء البنية قديمة وأيها ناشئة، وهذا غالبًا أنفع عرض مفرد لتحديد إلى أين يتحرك حقل.
Five concepts are enough to read any map produced by this method, and one of them is a warning about what the picture does not mean.
Nodes are the units. In a co-citation map they are cited works or authors; in co-word analysis they are terms. Node size in most visualisations represents frequency, so a large node is one that occurs often, not one that is important in any other sense.
Edges are relations, and each has a strength: how often the two nodes co-occur. Most software shows only edges above a threshold, so the visible connections are a subset chosen by a parameter you set.
Clusters are groups of nodes more connected to each other than to the rest, identified by an algorithm. The number of clusters is not a property of the field; it depends on a resolution parameter, and increasing it splits clusters into smaller ones. This is the most important thing to understand about the method, because a chapter reporting that a field has five clusters has reported a parameter choice as a finding.
Centrality measures how important a node is to the network's structure. Different centrality measures capture different kinds of importance: how many connections a node has, and how often it sits on the path between other nodes. The second identifies works that bridge clusters, which are frequently the most interesting items on a map, because a paper connecting two otherwise separate groups is doing integrative work.
Layout positions nodes so that connected ones sit near each other. The critical warning is that position has no meaning beyond proximity: there is no horizontal or vertical axis, no direction, and no scale. A node on the left is not more anything than a node on the right, and readers unfamiliar with network visualisation routinely over-interpret position.
One further caution about visual impression. Layout algorithms are partly stochastic, so running the same analysis twice can produce maps that look different while representing identical structure. Report the software and its settings, and do not treat the visual arrangement as a finding.
خمسة مفاهيم تكفي لقراءة أي خريطة يُنتجها هذا المنهج، وواحد منها تحذيرٌ عمّا لا تعنيه الصورة.
والعقد هي الوحدات. ففي خريطة استشهاد مشترك تكون أعمالًا أو مؤلفين مستشهَدًا بهم؛ وفي تحليل الكلمات المشتركة تكون مصطلحات. وحجم العقدة في معظم التصويرات يمثّل التكرار، فالعقدة الكبيرة هي التي تقع كثيرًا، لا التي تهم بأي معنى آخر.
والحواف علاقات، ولكلٍّ قوة: كم مرة تتزامن العقدتان. ومعظم البرمجيات تُظهر الحواف فوق عتبة فقط، فالصلات المرئية مجموعةٌ فرعية اختارها معاملٌ ضبطته أنت.
والعناقيد مجموعات عقد أشد اتصالًا ببعضها منها بالبقية، تحددها خوارزمية. وعددُ العناقيد ليس خاصيةً للحقل؛ بل يعتمد على معامل دقة، وزيادته تقسّم العناقيد إلى أصغر. وهذا أهم ما يُفهَم عن المنهج، لأن فصلًا يُبلغ بأن لحقلٍ خمسةَ عناقيد قد أبلغ بخيار معامل بوصفه نتيجة.
والمركزية تقيس كم عقدةٌ مهمة لبنية الشبكة. ومقاييس المركزية المختلفة تلتقط أنواعًا مختلفة من الأهمية: كم صلةً تملك عقدة، وكم مرة تجلس على المسار بين عقد أخرى. والثاني يحدد الأعمال التي تجسر العناقيد، وهي كثيرًا ما تكون أطرف العناصر على خريطة، لأن ورقةً تصل مجموعتين منفصلتين لولاها تؤدي عملًا تكامليًا.
والتخطيط يضع العقد بحيث تجلس المتصلة متقاربة. والتحذير الحاسم أن الموضع لا معنى له وراء التقارب: فلا محور أفقي ولا رأسي، ولا اتجاه، ولا مقياس. والعقدةُ على اليسار ليست أكثر أي شيء من العقدة على اليمين، والقرّاء غير المعتادين على تصوير الشبكات يفرطون روتينيًا في تفسير الموضع.
وتحذير آخر عن الانطباع البصري. فخوارزميات التخطيط عشوائية جزئيًا، فتشغيلُ التحليل نفسه مرتين قد يُنتج خرائط تبدو مختلفة وهي تمثّل بنيةً متطابقة. أبلغ بالبرمجية وإعداداتها، ولا تعامل الترتيب البصري نتيجةً.
Every map is produced by a set of parameter choices, and different choices produce different maps from the same data. This is the method's central methodological issue.
| Parameter | What it controls | Effect of raising it |
|---|---|---|
| Minimum occurrences | Which items enter the analysis | Fewer, more central nodes; a cleaner but narrower map |
| Minimum link strength | Which relations are counted | Sparser network; weak connections disappear |
| Clustering resolution | How finely the network is divided | More and smaller clusters |
| Number of items shown | Visual density of the map | Less clutter; peripheral items lost |
| Normalisation method | How co-occurrence is scaled | Changes which nodes appear central |
The consequence is that a map is not a fact about the field; it is a fact about the field under a stated set of parameters. This does not make the method subjective, but it does impose two obligations that many published analyses fail.
Report every parameter. The software, the version, each threshold, the clustering method and its resolution, and the normalisation. Without these the analysis cannot be reproduced, and a reader who runs the same corpus with default settings will get a different picture and will not know why.
Test sensitivity. Run the analysis at two or three parameter settings and see whether the structure holds. If the same broad clusters appear across settings, the structure is robust and you can report it with confidence. If the clusters rearrange completely when the resolution changes by a small amount, the structure is not stable and reporting one version as the field's structure is a misrepresentation.
Two practical rules help with the choices themselves. Set thresholds from the data, not from convention. Look at the distribution of occurrences and choose a threshold where there is a natural break, and explain that choice. A default of five because the software suggested it is a choice nobody made.
And prefer interpretability to completeness. A map with sixty labelled nodes that a reader can actually read is more useful than one with four hundred that is a coloured cloud. The purpose is communication, and an unreadable map communicates only that an analysis was run.
كل خريطة تُنتجها مجموعةُ خيارات معاملات، وخياراتٌ مختلفة تُنتج خرائط مختلفة من البيانات نفسها. وهذه المسألة المنهجية المركزية للمنهج.
| المعامل | ما يتحكم به | أثر رفعه |
|---|---|---|
| الحد الأدنى للتكرار | أي عناصر تدخل التحليل | عقد أقل وأكثر مركزية؛ خريطة أنظف وأضيق |
| الحد الأدنى لقوة الرابط | أي علاقات تُعَدّ | شبكة أرق؛ والصلات الضعيفة تختفي |
| دقة العنقدة | كم تُقسَّم الشبكة بدقة | عناقيد أكثر وأصغر |
| عدد العناصر المعروضة | الكثافة البصرية للخريطة | زحام أقل؛ وفقدان العناصر الطرفية |
| منهج التطبيع | كيف يُقاس التزامن | يغيّر أي العقد تظهر مركزية |
والنتيجة أن الخريطة ليست واقعةً عن الحقل؛ بل واقعةٌ عن الحقل تحت مجموعة معاملات مذكورة. وهذا لا يجعل المنهج ذاتيًا، لكنه يفرض التزامين تخفق فيهما تحليلات منشورة كثيرة.
أبلغ بكل معامل. البرمجية والإصدار وكل عتبة ومنهج العنقدة ودقته والتطبيع. فبدونها لا يمكن إعادة إنتاج التحليل، والقارئ الذي يشغّل المدونة نفسها بالإعدادات الافتراضية سيحصل على صورة مختلفة ولن يعرف لماذا.
واختبر الحساسية. شغّل التحليل بإعدادين أو ثلاثة وانظر هل تصمد البنية. فإن ظهرت العناقيد العريضة نفسها عبر الإعدادات، فالبنية متينة وتستطيع الإبلاغ بها بثقة. وإن أعادت العناقيد ترتيبها كليًا حين تتغير الدقة بقدر صغير، فالبنية غير مستقرة والإبلاغُ بنسخة واحدة بوصفها بنية الحقل تحريف.
وقاعدتان عمليتان تساعدان في الخيارات نفسها. اضبط العتبات من البيانات لا من العُرف. انظر في توزيع التكرارات واختر عتبةً عند انقطاع طبيعي، واشرح ذلك الاختيار. فافتراضُ خمسة لأن البرمجية اقترحته خيارٌ لم يتخذه أحد.
وفضّل القابلية للتفسير على الاكتمال. فخريطةٌ بستين عقدة مسمّاة يستطيع القارئ قراءتها فعلًا أنفع من واحدة بأربعمئة هي سحابة ملوّنة. فالغرض التواصل، والخريطةُ غير المقروءة توصّل فقط أن تحليلًا شُغِّل.
Producing a map takes an afternoon. Turning it into a contribution takes a week, and the week is the part that matters.
The procedure has four steps. Identify the clusters and, for each, list the most central items by whichever centrality measure suits. Read the central items, at least the five most central in each cluster, at the level of the abstract and introduction. Name each cluster from what you read rather than from the most frequent term, since a cluster's most frequent keyword is often a generic word that appears in everything.
Then describe the structure: what the clusters are, how they relate, which items bridge them, and what is absent. That last question is the one that most often produces a contribution, because a map makes absence visible in a way reading does not.
Four kinds of finding come out of this reliably. An unexpected cluster, meaning a group of work you had not treated as a distinct area, which tells you your framework was missing something. An unexpected separation, where two literatures you assumed were connected turn out to have almost no citation contact, which is frequently a real gap and an argument for a study that bridges them.
A bridging item, meaning a work with high betweenness that connects clusters, which is usually worth reading closely because it is doing integrative work the field values. And a temporal pattern, visible in an overlay by year, showing which parts of the structure are established and which are emerging, which directly supports a claim about where a field is moving.
Reporting has a fixed shape. The corpus as for any bibliometric analysis. The technique and its parameters in full. The map, at a size where labels are readable, with a caption stating what the nodes and edges represent. A table of clusters with the name you gave each, the number of items, and the most central items. And the prose, which is the substance: what each cluster is about, how they relate, and what follows for your study.
One closing observation about the method's place. Science mapping is at its best as the first step of a review rather than the whole of one. It tells you the field's structure and which papers are central, and that is precisely the information you need to decide what to read closely. Used that way it makes the reading better targeted and the resulting review more defensible, because the selection of papers to engage with came from the data rather than from what you happened to find first.
إنتاج خريطة يستغرق بعد ظهيرة. وتحويلها إلى إسهام يستغرق أسبوعًا، والأسبوع هو الجزء الذي يهم.
وللإجراء أربع خطوات. حدّد العناقيد ولكلٍّ اسرد أكثر العناصر مركزيةً بأي مقياس مركزية يلائم. واقرأ العناصر المركزية، خمسةً على الأقل من الأكثر مركزيةً في كل عنقود، بمستوى الملخص والمقدمة. وسمِّ كل عنقود ممّا قرأت لا من أكثر المصطلحات تكرارًا، لأن أكثر الكلمات المفتاحية تكرارًا في عنقود كثيرًا ما تكون كلمةً عامة تظهر في كل شيء.
ثم صِف البنية: ما العناقيد، وكيف ترتبط، وأي العناصر تجسرها، وما الغائب. وذلك السؤال الأخير هو الأكثر إنتاجًا لإسهام، لأن الخريطة تجعل الغياب مرئيًا بما لا تفعله القراءة.
وأربعة أنواع من النتائج تخرج من هذا بثبات. عنقودٌ غير متوقَّع، أي مجموعة أعمال لم تعاملها مجالًا متمايزًا، وهذا يخبرك أن إطارك كان يفتقد شيئًا. وانفصالٌ غير متوقَّع، حيث تتبيّن أدبيتان افترضت اتصالهما بلا تماس استشهادي تقريبًا، وهذا كثيرًا ما يكون فجوةً حقيقية وحجةً لدراسة تجسرهما.
وعنصرٌ جاسر، أي عمل بمركزية وساطة عالية يصل عناقيد، ويستحق عادةً قراءةً متأنية لأنه يؤدي عملًا تكامليًا يقدّره الحقل. ونمطٌ زمني، مرئي في تصوير طبقي بالسنة، يُظهر أي أجزاء البنية راسخة وأيها ناشئة، وهذا يسند مباشرةً ادّعاءً عن إلى أين يتحرك حقل.
وللإبلاغ شكلٌ ثابت. المدونة كما لأي تحليل ببليومتري. والتقنية ومعاملاتها كاملةً. والخريطة، بحجم تكون فيه اللصاقات مقروءة، بتعليق يذكر ماذا تمثّل العقد والحواف. وجدول عناقيد بالاسم الذي أعطيته لكل واحد وعدد العناصر وأكثرها مركزية. والنثر، وهو الجوهر: عمّ كل عنقود، وكيف ترتبط، وما يترتب لدراستك.
وملاحظة ختامية عن موضع المنهج. فرسم خرائط العلم في أفضل حالاته أول خطوة في مراجعة لا المراجعةَ كلها. فهو يخبرك ببنية الحقل وأي الأوراق مركزية، وتلك بالضبط المعلومة التي تحتاجها لتقرير ما تقرؤه بإمعان. ومستخدَمًا هكذا يجعل القراءة أفضل استهدافًا والمراجعةَ الناتجة أقبل للدفاع، لأن انتقاء الأوراق للتعامل معها جاء من البيانات لا مما صادف أن وجدته أولًا.