TechBooky AI Assistant
TechBooky AI Assistant
👋 Welcome to TechBooky AI Assistant

I can help with:
🔎 Tech News
🤖 AI Topics
💻 Gadgets
☁️ Cloud
✍️ Guest Posts
📢 Advertising
🔗 Backlinks
📩 Newsletter
  • AI Search
  • Cryptocurrency
  • Earnings
  • Enterprise
  • About TechBooky
  • Submit Article
  • Advertise With TechBooky
  • Contact Us
TechBooky
  • African
  • AI
  • Metaverse
  • Gadgets
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
  • African
  • AI
  • Metaverse
  • Gadgets
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
TechBooky
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
Home Artificial Intelligence

Wikipedia Launches New AI Accessibility Project

Akinola Ajibola by Akinola Ajibola
October 1, 2025
in Artificial Intelligence, Service news
Share on FacebookShare on Twitter
Share this story

Send it to someone who should read it.

f Facebook X X in LinkedIn wa WhatsApp tg Telegram @ Email

In Brief
  • Wikimedia Deutschland announced on Wednesday that a new database will enable AI models to access Wikipedia’s vast amount of knowledge.
  • This new AI-friendly database is now being added to Wikidata, which will facilitate the information’s assimilation by huge language models.
  • This approach which is also known as the Wikidata Embedding Project, uses a vector-based semantic search, which is a method that aids computers in comprehending the...

Wikimedia Deutschland announced on Wednesday that a new database will enable AI models to access Wikipedia’s vast amount of knowledge.

This new AI-friendly database is now being added to Wikidata, which will facilitate the information’s assimilation by huge language models. This approach which is also known as the Wikidata Embedding Project, uses a vector-based semantic search, which is a method that aids computers in comprehending the meaning and connections between words, to search through the over 120 million articles that now make up Wikipedia and its sister platforms which in turns transforms the items in Wikidata from clumsily formatted data into vectors that reflect the context and meaning surrounding the Wikidata entry, the Berlin-based team used a huge language model over the past year.

This information is best visualised in vectorised style as a network of dots and interwoven lines. According to Lydia Pintscher, portfolio lead for Wikidata, Adams would be linked to the word “human” and the names of his books which was reported by a news media firm.

The initiative increases the data’s accessibility to natural language enquiries from LLMs as well as new support for the Model Context Protocol (MCP), a standard that facilitates communication between AI systems and data sources.

Wikimedia’s German division took on the project with Jina.AI and DataStax, a neural search and firm, which is IBM-owned, a real-time training data provider.

For years, Wikidata has provided machine-readable data from Wikimedia properties; however, the tools that were available at the time were limited to keyword searches and the specialised query language SPARQL. The new system will be more compatible with retrieval-augmented generation (RAG) systems, which enable AI models to draw in outside data, allowing developers to base their models on information that has been validated by Wikipedia editors.

Important semantic context is also provided by the data’s structure. For example, searching the database for the word “scientist” will get lists of both Bell Labs scientists and well-known nuclear physicists. A Wikimedia-approved image depicting scientists at work, translations of the word “scientist” into other languages, and extrapolations to related terms like “researcher” and “scholar” are also included.

Also worth reading
AI Coding Is Creating A New Burden For Open-Source Maintainers Paystack’s AI Checkout Experiment Could Change How Nigerians Pay Online Apple’s OpenAI Lawsuit Turns The AI Hardware Race Into A Legal Fight Sam Altman And Elon Musk Are Now Fighting Over Space Data Centres Nigeria’s Probe Into Meta, X, Alphabet And AI Firms Puts Big Tech’s Media Power On Trial AI Spending Is Starting To Look Like An Inflation Story, Not Just A Tech Story

Toolforge makes the database available to the general public. For developers who are interested, Wikidata will also be holding a webinar on October 9.

The new initiative comes as AI researchers which continue to look for reliable data sources to help them refine their models. Even while the training systems themselves have advanced and are now frequently put together as intricate training environments rather than straightforward datasets, they still need carefully selected data in order to work well. While some may despise Wikipedia, its data is far more fact-oriented than catch-all datasets like the Common Crawl, which is a vast collection of web pages scraped from all over the internet. This is especially important for deployments that require high accuracy.

The drive for high-quality data may occasionally have severe repercussions for AI labs. In August, Anthropic proposed to pay $1.5 billion to resolve a lawsuit against a group of writers whose writings had been used as training materials. This would put an end to any accusations of misconduct.

Philippe Saadé, the project manager for Wikidata AI, had stressed in a news release that his project is not affiliated with any major AI labs or tech businesses. Saadé told reporters, “This Embedding Project launch demonstrates that powerful AI doesn’t have to be controlled by a handful of companies.” “It can be open, cooperative, and designed with everyone’s best interests in mind.”

The researchers had converted Wikidata’s structured data, which was collected up until September 18, 2024, into vectors using a model from the AI company Jina AI. The infrastructure for storing the vector database for the project is presently provided for free by the IBM firm DataStax.

Before upgrading the database with data added over the past year, the team is awaiting input from developers who use it. Small changes or adjustments to already-existing Wikidata won’t make the database any less valuable, according to Saadé, even though the current database does not contain completely fresh material that has been uploaded in the last year. “The vector that we’re computing is ultimately just a general idea of an item, so even a minor edit made on Wikidata won’t have a significant impact,” he stated.

Related Reading

More contextual TechBooky stories selected from tags, categories and article context.

  • unnamed (66)
    Elon Musk Launches Grokipedia to Challenge Wikipedia
  • multiple-ai-chatbots-increasingly-cite-elon-musk-s
    AI Chatbots Increasingly Cite Musk’s Grokipedia…
  • Google-AI-Mode
    Google Brings African Languages Support To AI Search
  • assets_task_01jrw67y2ge32rms2shtmth089_img_0
    AI Search Engine Powered by OpenAI; Netflix Begins Testing
  • DataHub
    DataHub Turns SQL Query History into Context Layer…
  • meta
    Meta is Developing Its Own AI-Powered Search Engine
  • google-io-2023-051023-88
    Google Can Train Search AI on Content Without…
  • FILE PHOTO: OpenAI and ChatGPT logos are seen in this illustration taken, February 3, 2023. REUTERS/Dado Ruvic/Illustration/
    Why ChatGPT Has Sparked Unprecedented Interest
Keep Reading Smarter

Search TechBooky with AI

Use TechBooky's AI Search to explore the context behind this story and related coverage across the site.

Try AI Search
More On This Topic
Artificial Intelligence Service news
Follow TechBooky

Follow TechBooky for more technology stories and newsroom updates.

f Facebook X X in LinkedIn ig Instagram wa WhatsApp

Tags: AIwikimediawikipedia
Akinola Ajibola

Akinola Ajibola

Search TechBooky
Open TechBooky AI Search Try the AI Assistant

BROWSE BY CATEGORIES

Receive top tech news directly in your inbox

subscription from
Loading

Freshly Squeezed

  • SpaceX Deploys V3 Starlink Satellites But Loses Another Super Heavy Booster July 26, 2026
  • The Boring Company Reportedly Seeks $4B As Musk’s Tunnel Bet Gets A $20B Valuation July 26, 2026
  • Claude Opus 5 Gives Anthropic A Cheaper Answer To The Fable 5 Problem July 25, 2026
  • Meta Makes Facebook Verified Free As AI Scams Make Real People Harder To Spot July 24, 2026
  • SAP Cloud Growth Eases Fears That AI Will Weaken Enterprise Software July 24, 2026
  • US Lawmakers Push AI Kill Switch Bill After OpenAI Rogue-Model Incident July 24, 2026
  • Airtel Money’s $61B Quarter Makes Its London IPO A Bigger Africa Fintech Story July 24, 2026
  • Intel Q2 Revenue Jumps As AI Compute Demand Lifts Chip Business July 24, 2026
  • AMD And Anthropic Deal Puts Real Pressure On Nvidia’s AI Chip Lead July 24, 2026
  • New York’s Data Centre Pause Shows AI Infrastructure Is Hitting Politics July 23, 2026
  • OpenAI Researcher’s $2B Drug Discovery Plan Shows AI Biotech Hype Is Back July 23, 2026
  • Airtel Africa Q1 Shows Mobile Money And Data Are Doing The Heavy Lifting July 23, 2026

Browse Archives

July 2026
M T W T F S S
 12345
6789101112
13141516171819
20212223242526
2728293031  
« Jun    

Quick Links

  • About TechBooky
  • Advertise With TechBooky
  • Contact us
  • Submit Article
  • Privacy Policy
Generic selectors
Exact matches only
Search in title
Search in content
Post Type Selectors
Search in posts
Search in pages
  • African
  • Artificial Intelligence
  • Gadgets
  • Metaverse
  • Tips
  • AI Search
  • About TechBooky
  • Advertise With TechBooky
  • Submit Article
  • Contact us

© 2025 Designed By TechBooky Elite

Discover more from TechBooky

Subscribe now to keep reading and get access to the full archive.

Continue reading

We use cookies to ensure that we give you the best experience on our website. If you continue to use this site we will assume that you are happy with it.