MLKN.lab · Method
Method
How we build, classify, and model the knowledge networks.
Overview
A Polyhierarchical, Data-Driven Approach
The operational architecture of MLKN.lab relies on a formalized, multi-layered data ingestion and network-synthesis pipeline designed to map the global topology of science. Moving away from rigid, mono-hierarchical taxonomies, our methodology leverages a high-fidelity, polyhierarchical design that accommodates the cross-disciplinary fluidity of modern discovery.
By ingesting large-scale bibliometric data from OpenAlex, the pipeline normalizes and maps complex scientific metadata across five discrete ontological layers. Through co-occurrence analysis, topological graph algorithms, and rigorous modularity validation, MLKN.lab extracts structural insights from chaotic data, rendering a reproducible, open-access, and interactive geometric representation of human knowledge.
MLKN-lab’s methodology combines bibliometrics, network science, and hierarchical classification to map the structure of scientific knowledge. Our approach is: data-driven, reproductible, open, and interdiscilplinar.
Methodological Foundations
Network Science, Multilayer Networks, Hypergraphs, Semantic Web, Collective Intelligence
Network Science
Applying graph theory and network analysis to model scientific knowledge. MLKN.lab represents disciplines, subdisciplines, and concepts as nodes and edges, enabling the study of their topological properties.
Key Tools:
- Graph Centrality: Identifying key nodes in knowledge networks.
- Community Detection: Finding clusters of related concepts.
- Path Analysis: Tracing connections between distant fields.
Multilayer Networks
Modeling scientific knowledge as multilayer networks, where each layer represents a distinct dimension (e.g., disciplines, time, or scale). MLKN.lab uses multilayer network analysis to study how knowledge evolves across layers.
Key Tools:
- Layer Interdependencies: Analyzing connections between layers.
- Cross-Layer Dynamics: Studying how changes in one layer affect others.
- Structural Controllability: Identifying key nodes that influence the entire network.
Hypergraphs
Representing scientific knowledge as hypergraphs, where edges can connect any number of nodes. MLKN.lab uses hypergraphs to model polyhierarchical relationships in knowledge systems.
Key Tools:
- Hypergraph Neural Networks: Applying deep learning to hypergraphs.
- Structure Preservation: Maintaining relationships in high-dimensional data.
- Topological Analysis: Studying the geometric properties of hypergraphs.
Semantic Web
Using semantic technologies to model the meaning of scientific concepts. MLKN.lab integrates ontologies, taxonomies, and knowledge graphs to enable semantic reasoning over scientific knowledge.
Key Tools:
- RDF/OWL: Formal representations of knowledge.
- Linked Data: Connecting knowledge across the web.
- SPARQL: Querying semantic knowledge bases.
Collective Intelligence
Studying how knowledge emerges from communities. MLKN.lab explores collaborative problem-solving, swarm intelligence, and crowd wisdom to understand how groups create and diffuse knowledge.
Key Tools:
- Social Network Analysis: Mapping collaborations.
- Cognitive Modeling: Simulating group decision-making.
- Swarm Algorithms: Optimizing collective behavior.
Tools and Techniques
OpenAlex, ORKG, Wikidata
OpenAlex
A free, open catalog of 200M+ scholarly works, including papers, authors, institutions, and concepts. MLKN.lab uses OpenAlex as its primary data source for building its polyhierarchical knowledge hypergraph.
Key Features:
- Comprehensive Metadata: Titles, authors, abstracts, citations, and more.
- Interconnected Data: Relationships between papers, authors, and institutions.
- Open Access: Free to use with no restrictions.
Open Research Knowledge Graph (ORKG)
A collaborative, open knowledge graph that represents research contributions as interconnected nodes. MLKN.lab aligns with ORKG’s mission to compare, analyze, and discover research across disciplines.
Key Features:
- Research Contributions: Papers, datasets, software, and more.
- Comparative Analysis: Compare research across fields.
- Open Data: Free and open-access.
Wikidata
A free, open knowledge graph that connects data from Wikipedia and other Wikimedia projects. MLKN.lab uses Wikidata to integrate and analyze diverse datasets.
Key Features:
- Structured Data: Millions of interconnected concepts.
- Open Access: Free and open for all.
- Linked Data: Connects to the broader semantic web.
OpenAlex: A Strategic Choice
Explanation of our strategic choice for OpenAlex
Why OpenAlex?
At MLKN.lab, we believe that knowledge should be open, accessible, and sovereign. That’s why we built our polyhierarchical knowledge hypergraph on OpenAlex, a free, open catalog of scholarly papers that democratizes access to research metadata.
OpenAlex is more than a database—it’s a movement. By providing comprehensive, interconnected, and up-to-date data on publications, authors, and institutions, OpenAlex enables tools like MLKN.lab to map, analyze, and simulate the structure of scientific knowledge without relying on commercial, restrictive databases.
We’re thrilled to see institutions like the CNRS—one of the world’s leading research organizations— transitioning to OpenAlex as part of their commitment to open science and research sovereignty. This decision validates our approach and signals a new era for research—one where knowledge is free, transparent, and collaborative.
Data Collection
Sources and Preprocessing
Primary Data Source: OpenAlex
We use OpenAlex, a free and open catalog of scholarly papers, authors, venues, and institutions. OpenAlex provides:
- Over 200M works (papers, preprints, etc.).
- Metadata (titles, abstracts, authors, venues, citations).
- Concepts and topics extracted from papers.
Supplementary Data Sources
To enrich our dataset, we supplement OpenAlex with:
- Scopus: For citation and thematic data.
- MeSH: For biomedical classifications.
- IEEE Thesaurus: For engineering and computer science.
Data Preprocessing
Raw data is cleaned and standardized to ensure consistency:
- Deduplication: Removing duplicate entries.
- Normalization: Standardizing names, terms, and classifications.
- Disambiguation: Resolving ambiguous author or venue names.
Classification
Organizing Knowledge into Hierarchies
Ontological Layers
We organize knowledge into 5 ontological layers:
- Core Discipline Domains: 6 high-level domains (e.g., Natural Sciences, Social Sciences).
- Academic Disciplines: 25 fields (e.g., Psychology, Computer Science).
- Subdisciplines: 235 specialized areas (e.g., Cognitive Psychology, AI Ethics).
- Core Thematic Domains: Thematic groupings within subdisciplines.
- Main Concepts: Specific topics or ideas (e.g., "Attention", "Neural Networks").
Classification Standards
Our hierarchy aligns with global standards:
- OECD Frascati Manual: For discipline classifications.
- UNESCO Fields of Science: For broad domain groupings.
- MeSH and IEEE Thesaurus: For specialized fields.
Polyhierarchical Design
Unlike traditional hierarchies, our structure is polyhierarchical:
- A concept can belong to multiple disciplines (e.g., AI Ethics → Computer Science + Philosophy).
- Disciplines can be grouped under multiple domains (e.g., Neuroscience → Natural Sciences + Health Sciences).
- This reflects the interdisciplinary nature of modern science.
Network Construction
Mapping Connections Between Disciplines
Co-Occurrence Analysis
We identify connections between disciplines by analyzing:
- Shared concepts: Topics that appear in multiple disciplines.
- Citation networks: How often papers from one discipline cite another.
- Author collaborations: Researchers working across disciplines.
Network Metrics
We use graph theory to quantify the structure of knowledge:
- Centrality: Identifying key disciplines and concepts.
- Modularity: Detecting communities or clusters.
- Bridges: Mapping connections between fields.
Visualization
Networks are visualized using interactive tools:
- Force-directed layouts: For exploring connections.
- Hierarchical views: For navigating layers.
- Color-coding: By domain, discipline, or subdiscipline.
Validation
Ensuring Accuracy and Relevance
Expert Review
Our hierarchy and networks are reviewed by domain experts to ensure:
- Accuracy: Disciplines and subdisciplines are correctly classified.
- Relevance: Connections reflect real-world interdisciplinary links.
- Completeness: No major fields or connections are missing.
Network Metrics
We validate the structure using:
- Modularity scores: To assess the strength of disciplinary clusters.
- Centrality measures: To identify key disciplines.
- Path analysis: To trace connections between fields.
Reproducibility
All our methods and datasets are open and reproducible:
- Code: Available on GitHub.
- Data: Publicly accessible datasets.
- Documentation: Detailed methodologies and tutorials.
Tools and Code
Software for Building and Analyzing Knowledge Networks
Data Processing Pipeline
A Python-based pipeline for cleaning, classifying, and structuring raw data from OpenAlex and other sources.
Network Analysis Scripts
Scripts for analyzing knowledge networks, including centrality metrics, community detection, and visualization.
Interactive Visualization Tools
Tools for exploring knowledge networks, including force-directed layouts and hierarchical views.
Explore Our Methodology
Dive deeper into our data, code, and tools.