VOSviewer is free software for building and viewing bibliometric networks. It is the most widely used tool for this in business and management research, it runs on every platform, and a working map takes about twenty minutes once you have the data.
Its scope is worth stating so expectations are right. It builds and visualises networks from bibliographic data, and it does that well. It does not perform performance analysis, statistical testing, or text analysis beyond term extraction, and it does not clean your data for you. For counting indicators and time series you need a different tool, and many researchers use a scripting package alongside it for exactly that reason.
The workflow is short: export records from a citation index, create a map by choosing a data type and an analysis, set thresholds, and interpret. Everything difficult is in the first and last steps, which are the export you got wrong and the interpretation you have not done, and both are matters of research design rather than of software.
This framework covers the workflow step by step, the data formats and their differences, the three views and what each shows, thesaurus files for cleaning, the settings that matter, and what to report so the analysis is reproducible. It is my own synthesis, written in my own words and grounded in recognized scholarship.
فوس فيور برمجية مجانية لبناء الشبكات الببليومترية وعرضها. وهي الأداة الأوسع استعمالًا لهذا في بحوث الأعمال والإدارة، وتعمل على كل منصة، وخريطةٌ عاملة تستغرق نحو عشرين دقيقة متى ملكت البيانات.
ونطاقها يستحق الذكر لتكون التوقعات صحيحة. فهي تبني الشبكات وتصوّرها من بيانات ببليوغرافية، وتفعل ذلك جيدًا. ولا تُجري تحليل أداء ولا اختبارًا إحصائيًا ولا تحليل نصوص وراء استخراج المصطلحات، ولا تنظّف بياناتك نيابةً عنك. ولمؤشرات العدّ والسلاسل الزمنية تحتاج أداةً أخرى، وباحثون كثيرون يستخدمون حزمة برمجة إلى جانبها لذلك السبب بالضبط.
وسير العمل قصير: صدّر سجلات من فهرس استشهاد، وأنشئ خريطة باختيار نوع بيانات وتحليل، واضبط العتبات، وفسّر. وكل ما هو صعب في الخطوتين الأولى والأخيرة، وهما التصدير الذي أخطأته والتفسير الذي لم تُنجزه، وكلتاهما مسألتا تصميم بحثي لا مسألتا برمجية.
ويغطي هذا الإطار سير العمل خطوةً خطوة، وصيغ البيانات وفروقها، والعروض الثلاثة وما يُظهره كلٌّ، وملفات المكنز للتنظيف، والإعدادات التي تهم، وما يُبلَّغ به ليكون التحليل قابلًا لإعادة الإنتاج. وقد أعددتُ هذا الإطار بنفسي وكتبتُه بأسلوبي، معتمدًا على المراجع العلمية المعتمدة.
Four steps produce a map. The first determines what is possible and the last determines whether it means anything.
Export. The software reads the export formats of the major citation indexes directly, and this is where most first attempts fail. The critical requirement is that the export must include cited references, which is usually an option you have to select rather than a default. Without them, citation-based analyses are unavailable and only co-word and co-authorship will run.
Two practical constraints apply. Indexes limit how many records can be exported at once, so a corpus of several thousand comes in batches, and the software accepts multiple files as one dataset. And the export must be the full record format rather than a brief one, since brief exports omit keywords and abstracts.
Create the map. You choose three things in sequence. The data type, which is bibliographic data files for the standard workflow. The analysis, which is one of the four relations: co-citation, bibliographic coupling, co-word, or co-authorship. And the unit of analysis, meaning documents, sources, authors, organizations, countries, or terms.
Then the counting method, which is a choice most people accept as a default without knowing what it does. Full counting gives each co-occurrence equal weight. Fractional counting divides the weight of a document among its links, which reduces the influence of documents with very long reference lists or many authors. Fractional counting is generally preferable for co-authorship and for corpora containing review articles, and the choice must be reported either way.
Set thresholds. The software asks for a minimum number of occurrences or citations and then shows how many items meet it, and you choose how many of those to keep. Both numbers must be reported, and both should be chosen from the distribution rather than accepted as defaults.
Clean and read. Apply a thesaurus file, regenerate, and then do the interpretive work. This last step is the analysis; the first three are data preparation.
أربع خطوات تُنتج خريطة. والأولى تحدد ما هو ممكن والأخيرة تحدد هل يعني شيئًا.
التصدير. البرمجية تقرأ صيغ تصدير فهارس الاستشهاد الكبرى مباشرةً، وهنا تخفق معظم المحاولات الأولى. والمتطلب الحاسم أن يشمل التصدير المراجع المستشهَد بها، وهذا عادةً خيارٌ عليك انتقاؤه لا افتراضي. وبدونها تكون التحليلات القائمة على الاستشهاد غير متاحة ولن يعمل إلا الكلمات المشتركة والتأليف المشترك.
وقيدان عمليان ينطبقان. فالفهارس تحدّ كم سجلًا يمكن تصديره دفعةً، فمدونةٌ من عدة آلاف تأتي دفعات، والبرمجية تقبل ملفات متعددة مجموعةَ بيانات واحدة. والتصدير يجب أن يكون صيغة السجل الكامل لا موجزةً، لأن التصديرات الموجزة تُغفل الكلمات المفتاحية والملخصات.
إنشاء الخريطة. تختار ثلاثة أشياء بالتتابع. نوع البيانات، وهو ملفات بيانات ببليوغرافية لسير العمل المعياري. والتحليل، وهو واحدة من العلاقات الأربع: الاستشهاد المشترك أو الاقتران الببليوغرافي أو الكلمات المشتركة أو التأليف المشترك. ووحدة التحليل، أي وثائق أو مصادر أو مؤلفين أو منظمات أو بلدان أو مصطلحات.
ثم منهج العدّ، وهو خيار يقبله معظم الناس افتراضيًا دون معرفة ما يفعله. فـالعدّ الكامل يعطي كل تزامن وزنًا متساويًا. والعدّ الكسري يقسّم وزن وثيقة على روابطها، وهذا يخفّض تأثير الوثائق ذات قوائم المراجع الطويلة جدًا أو المؤلفين الكثيرين. والعدّ الكسري مفضَّل عمومًا للتأليف المشترك وللمدونات التي تحوي مقالات مراجعة، والخيار يجب الإبلاغ به في الحالتين.
ضبط العتبات. تسأل البرمجية عن حد أدنى لعدد التكرارات أو الاستشهادات ثم تُظهر كم عنصرًا يستوفيه، وتختار أنت كم منها تُبقي. والرقمان يجب الإبلاغ بهما، وكلاهما ينبغي اختياره من التوزيع لا قبوله افتراضيًا.
التنظيف والقراءة. طبّق ملف مكنز، وأعد التوليد، ثم أنجز العمل التفسيري. وهذه الخطوة الأخيرة هي التحليل؛ والثلاث الأولى إعدادُ بيانات.
One map, three ways of looking at it. Each answers a different question, and one of them is consistently underused.
The network view shows nodes coloured by cluster, sized by frequency, and positioned so that connected items sit close together. This is the standard image reproduced in papers and it answers the question of what the field's structure is. Its limitation is that with more than about a hundred labelled items it becomes unreadable, which is why threshold choice matters as much for communication as for analysis.
The overlay view colours nodes by a variable, most usefully the average publication year of the documents in which each item appears. This produces a picture of the field's development: older parts of the structure in one colour, emerging parts in another. For a doctoral literature review this is frequently the single most useful output of the entire method, because it identifies which parts of a field are established and which are current, and that distinction directly supports a gap claim.
The overlay can also be driven by other variables, including average citation counts, which shows which parts of the structure are highly cited and which are not, and which is a different question from which parts are recent.
The density view replaces the nodes with a heat map showing where items concentrate. It is useful for a quick view of a large map where labels are unreadable, and it hides the individual items, so it is a supplement rather than a substitute.
Two practical points about producing images for a thesis. Export at high resolution and check that labels are legible at print size, since the default screen export is usually too small. And consider showing two views of the same map, typically network and overlay, because they answer different questions and together they make an argument that either alone does not.
A note on colour. Cluster colours are assigned arbitrarily and carry no meaning across maps, so cluster one in one analysis has no relationship to cluster one in another. Say this in the caption if there is any chance of confusion, and check that the colours remain distinguishable in greyscale for the printed version.
خريطة واحدة، وثلاث طرق للنظر إليها. وكلٌّ تجيب سؤالًا مختلفًا، وواحدةٌ منها قليلة الاستعمال بثبات.
وعرض الشبكة يُظهر العقد ملوَّنة بالعنقود، ومحجَّمة بالتكرار، وموضوعةً بحيث تجلس العناصر المتصلة متقاربة. وهذه الصورة المعيارية المنقولة في الأوراق وهي تجيب سؤال ما بنية الحقل. وحدّها أنه بأكثر من نحو مئة عنصر مُسمّى تصير غير مقروءة، ولهذا يهم اختيار العتبة للتواصل بقدر ما يهم للتحليل.
والعرض الطبقي يلوّن العقد بمتغير، وأنفعه متوسط سنة نشر الوثائق التي يظهر فيها كل عنصر. وهذا يُنتج صورةً لتطور الحقل: أجزاء البنية الأقدم بلون، والناشئة بآخر. ولمراجعة أدبيات في الدكتوراه يكون هذا كثيرًا أنفع مخرَج مفرد للمنهج كله، لأنه يحدد أي أجزاء الحقل راسخة وأيها حالية، وذلك التمييز يسند مباشرةً ادّعاء فجوة.
والعرض الطبقي يمكن أن يقوده متغيرات أخرى أيضًا، بما فيها متوسط أعداد الاستشهاد، وهذا يُظهر أي أجزاء البنية كثيرة الاستشهاد وأيها ليست كذلك، وهذا سؤال غير سؤال أي الأجزاء حديثة.
وعرض الكثافة يستبدل بالعقد خريطةً حرارية تُظهر أين تتركّز العناصر. وهو نافع لعرض سريع لخريطة كبيرة تكون اللصاقات فيها غير مقروءة، وهو يخفي العناصر المفردة، فهو مكمّل لا بديل.
ونقطتان عمليتان عن إنتاج الصور لأطروحة. صدّر بدقة عالية وافحص أن اللصاقات مقروءة بحجم الطباعة، فالتصدير الافتراضي للشاشة صغير جدًا عادةً. وفكّر في عرض عرضين للخريطة نفسها، شبكةً وطبقيًا عادةً، لأنهما يجيبان سؤالين مختلفين ومعًا يصنعان حجةً لا يصنعها أيٌّ منهما وحده.
وملاحظة عن اللون. ألوان العناقيد تُسنَد اعتباطيًا ولا تحمل معنى عبر الخرائط، فالعنقود الأول في تحليل لا علاقة له بالعنقود الأول في آخر. قُل هذا في التعليق إن كان ثمة احتمال التباس، وافحص أن الألوان تبقى متمايزة بتدرج رمادي للنسخة المطبوعة.
Cleaning is what separates a meaningful map from a picture of your data's inconsistencies. The mechanism is a thesaurus file, and building one is not optional.
| Problem | Example | Effect if uncleaned |
|---|---|---|
| Keyword synonyms | SME, SMEs, small and medium enterprise | One theme split across three nodes |
| Singular and plural | capability, capabilities | Both below threshold; the theme disappears |
| Spelling variants | organisation, organization | The field's central term fragmented |
| Author name forms | Smith J, Smith John, J Smith | One researcher's output split three ways |
| Institution variants | renamed or merged organizations | Misleading institutional rankings |
| Generic terms | study, research, paper, analysis | Meaningless hubs dominating the map |
A thesaurus file is a two-column list: the variant, and the label it should be replaced with. Leaving the replacement blank removes the item entirely, which is how you drop generic terms. The file is loaded when the map is created, and the same file can be reused across analyses of the same corpus.
Building one is a bounded task. Create the map without a thesaurus, export the item list, sort it, and work down it merging variants and marking terms for removal. For a corpus of a few thousand records this takes two to four hours and it is the highest-value time in the whole analysis, because every result downstream depends on it.
Three specific decisions arise. Generic terms such as study, research, results, and analysis appear in almost every record and produce large central nodes that mean nothing. Remove them, and say in the reporting that you did.
Method terms such as survey, case study, and regression are a judgement call: in a map of a field's topics they are noise, and in a map intended to show methodological structure they are the point. Decide according to your question and state the decision.
And the unit of the concept: whether to merge a broad term with its narrower variants depends on whether the distinction matters for your question. Merging digital transformation with digitalisation may be right for a broad map and wrong for a study distinguishing them, and this is exactly the kind of decision that must be reported rather than made silently.
التنظيف هو ما يفصل خريطةً ذات معنى عن صورةٍ لعدم اتساق بياناتك. والآلية ملفُ مكنز، وبناؤه ليس اختياريًا.
| المشكلة | مثال | الأثر إن لم يُنظَّف |
|---|---|---|
| مترادفات الكلمات المفتاحية | المنشأة الصغيرة والمتوسطة ومختصرها وجمعها | موضوع واحد مقسَّم على ثلاث عقد |
| المفرد والجمع | قدرة، قدرات | كلاهما تحت العتبة؛ فيختفي الموضوع |
| تنويعات الهجاء | organisation وorganization | مصطلح الحقل المركزي مجزَّأ |
| صيغ أسماء المؤلفين | ثلاث صيغ للاسم نفسه | إنتاج باحث واحد مقسَّم ثلاث مرات |
| تنويعات المؤسسات | منظمات أُعيدت تسميتها أو اندمجت | ترتيبات مؤسسية مضللة |
| المصطلحات العامة | دراسة، بحث، ورقة، تحليل | محاور بلا معنى تهيمن على الخريطة |
وملف المكنز قائمةٌ بعمودين: التنويعة، واللصاقة التي ينبغي استبدالها بها. وتركُ الاستبدال فارغًا يزيل العنصر كليًا، وهكذا تُسقط المصطلحات العامة. ويُحمَّل الملف عند إنشاء الخريطة، ويمكن إعادة استخدام الملف نفسه عبر تحليلات المدونة نفسها.
وبناؤه مهمةٌ محدودة. أنشئ الخريطة بلا مكنز، وصدّر قائمة العناصر، وافرزها، واعمل عليها نازلًا دامجًا التنويعات ومعلّمًا المصطلحات للإزالة. ولمدونة من بضعة آلاف سجل يستغرق هذا ساعتين إلى أربع وهو أعلى وقت قيمةً في التحليل كله، لأن كل نتيجة في المصبّ تعتمد عليه.
وثلاثة قرارات محددة تنشأ. المصطلحات العامة مثل «دراسة» و«بحث» و«نتائج» و«تحليل» تظهر في كل سجل تقريبًا وتُنتج عقدًا مركزية كبيرة لا تعني شيئًا. أزلها، وقُل في الإبلاغ أنك فعلت.
ومصطلحات المنهج مثل «مسح» و«دراسة حالة» و«انحدار» حكمٌ: ففي خريطة موضوعات حقل تكون ضوضاءً، وفي خريطة يُقصَد بها إظهار البنية المنهجية تكون هي المقصود. قرّر بحسب سؤالك واذكر القرار.
ووحدة المفهوم: فهل تدمج مصطلحًا عامًا بتنويعاته الأضيق يعتمد على هل يهم التمييز لسؤالك. فدمجُ «التحول الرقمي» بـ«الرقمنة» قد يصح لخريطة عريضة ويخطئ لدراسة تميّز بينهما، وهذا بالضبط نوع القرار الذي يجب الإبلاغ به لا اتخاذه صمتًا.
A mapping analysis is reproducible only if the parameters are reported. Eight items cover it and they take a paragraph.
The corpus: database, search string, fields, filters, date of extraction, and the record count. The export: format and whether cited references were included. The software: name and version. The analysis: which relation, which unit of analysis, and the counting method.
The thresholds: minimum occurrences or citations, how many items met the threshold, and how many were retained. The cleaning: that a thesaurus was used, how many merges it contained, and what categories of term were removed. The clustering: the resolution and any minimum cluster size. And the interpretation basis: which items you read in order to name each cluster.
That last item is the one that distinguishes a research contribution from a software output, and it is almost never reported. Saying that each cluster was named after reading its five most central documents tells a reader that interpretation happened, and its absence tells them it may not have.
Five problems recur in practice and each has a fixed cause. An empty or tiny map means the threshold is too high or the export lacked cited references; check the export first. An unreadable map means too many items; raise the threshold rather than shrinking the image.
One enormous central node is almost always a generic term or an unmerged variant, and the thesaurus fixes it. Clusters that make no sense usually mean the corpus is too heterogeneous, because a search string that caught two unrelated literatures produces two unrelated clusters, which is a corpus problem rather than an analysis problem.
And results that change between runs are normal for the layout, which is partly stochastic, and are not normal for the clustering; if cluster membership changes substantially, the structure is unstable and that should be reported rather than resolved by picking a run.
A closing point about the tool's place in a thesis. VOSviewer produces a picture in twenty minutes, and that speed is exactly why it is misused: a chapter can be filled with maps without any of the interpretive work that makes them mean something. The discipline is to treat each map as a question rather than an answer, to read the central items in each cluster, and to write the prose that says what the structure is and what follows from it for your study. The software does the drawing; the contribution is entirely in the reading.
تحليل الرسم قابل لإعادة الإنتاج فقط إن أُبلغ بالمعاملات. وثمانية بنود تغطيه وتستغرق فقرة.
المدونة: القاعدة وسلسلة البحث والحقول والمرشِّحات وتاريخ الاستخراج وعدد السجلات. والتصدير: الصيغة وهل شملت المراجع المستشهَد بها. والبرمجية: الاسم والإصدار. والتحليل: أي علاقة وأي وحدة تحليل ومنهج العدّ.
والعتبات: الحد الأدنى للتكرارات أو الاستشهادات، وكم عنصرًا استوفى العتبة، وكم أُبقي. والتنظيف: أن مكنزًا استُخدم، وكم دمجًا حواه، وأي فئات مصطلحات أُزيلت. والعنقدة: الدقة وأي حد أدنى لحجم العنقود. وأساس التفسير: أي العناصر قرأتَ لتسمية كل عنقود.
وذلك البند الأخير هو ما يميّز إسهامًا بحثيًا عن مخرَج برمجية، ولا يكاد يُبلَّغ به أبدًا. فقولُ إن كل عنقود سُمِّي بعد قراءة أكثر خمس وثائق فيه مركزيةً يخبر القارئ أن تفسيرًا وقع، وغيابُه يخبره أنه قد لا يكون وقع.
وخمس مشكلات تتكرر عمليًا ولكلٍّ سببٌ ثابت. خريطةٌ فارغة أو ضئيلة تعني أن العتبة عالية جدًا أو أن التصدير افتقر إلى المراجع المستشهَد بها؛ افحص التصدير أولًا. وخريطةٌ غير مقروءة تعني عناصر كثيرة جدًا؛ ارفع العتبة لا تصغّر الصورة.
وعقدةٌ مركزية هائلة واحدة شبه دائمًا مصطلحٌ عام أو تنويعةٌ غير مدموجة، والمكنز يصلحها. وعناقيد لا معنى لها تعني عادةً أن المدونة مفرطة التباين، لأن سلسلة بحث التقطت أدبيتين غير مترابطتين تُنتج عنقودين غير مترابطين، وهذه مشكلة مدونة لا مشكلة تحليل.
ونتائج تتغير بين التشغيلات طبيعيةٌ للتخطيط، وهو عشوائي جزئيًا، وليست طبيعية للعنقدة؛ فإن تغيّرت عضوية العناقيد كثيرًا، فالبنية غير مستقرة وينبغي الإبلاغ بذلك لا حلّه بانتقاء تشغيلة.
ونقطة ختامية عن موضع الأداة في أطروحة. فوس فيور يُنتج صورةً في عشرين دقيقة، وتلك السرعة بالضبط لماذا يُساء استخدامه: فيمكن ملء فصل بالخرائط بلا أي من العمل التفسيري الذي يجعلها تعني شيئًا. والانضباط معاملةُ كل خريطة سؤالًا لا إجابة، وقراءةُ العناصر المركزية في كل عنقود، وكتابةُ النثر الذي يقول ما البنية وما يترتب عليها لدراستك. فالبرمجية تؤدي الرسم؛ والإسهام كله في القراءة.