Loading...
Thumbnail Image
Publication

Exploring food contents in scientific literature with FoodMine

Editors
Title / Series / Name
Scientific Reports
Publication Volume
10
Publication Issue
1
Pages
Editors
Keywords
URI
http://hdl.handle.net/20.500.14018/13895
Abstract
Thanks to the many chemical and nutritional components it carries, diet critically affects human health. However, the currently available comprehensive databases on food composition cover only a tiny fraction of the total number of chemicals present in our food, focusing on the nutritional components essential for our health. Indeed, thousands of other molecules, many of which have well documented health implications, remain untracked. To explore the body of knowledge available on food composition, we built FoodMine, an algorithm that uses natural language processing to identify papers from PubMed that potentially report on the chemical composition of garlic and cocoa. After extracting from each paper information on the reported quantities of chemicals, we find that the scientific literature carries extensive information on the detailed chemical components of food that is currently not integrated in databases. Finally, we use unsupervised machine learning to create chemical embeddings, finding that the chemicals identified by FoodMine tend to have direct health relevance, reflecting the scientific community’s focus on health-related chemicals in our food.
Topic
Publisher
Place of Publication
Type
Journal article
Date
2020
Language
ISBN
Identifiers
10.1038/s41598-020-73105-0
Publisher link
Unit