TY - GEN
T1 - PayDIBI: Pay-as-you-go data integration for bioinformatics
AU - Wanders, B.
PY - 2012/12/10
Y1 - 2012/12/10
N2 - Background: Scientific research in bio-informatics is often data-driven and supported by biolog-
ical databases. In a growing number of research projects, researchers like to ask questions that
require the combination of information from more than one database. Most bio-informatics papers
do not detail the integration of different databases. As roughly 30% of all tasks in workflows are
data transformation tasks, database integration is an important issue.
Integrating multiple data sources can be difficult. As data sources are created, many design
decisions are made by their creators.
Methods: Our research is guided by two use cases: homologues, the representation and integration
of groupings; metabolomics integration, with a focus on the TCA cycle.
Results: We propose to approach the time consuming problem of integrating multiple biological
databases through the principles of ‘pay-as-you-go’ and ‘good-is-good-enough’. By assisting the
user in defining a knowledge base of data mapping rules, trust information and other evidence we
allow the user to focus on the work, and put in as little effort as is necessary for the integration.
Through user feedback on query results and trust assessments, the integration can be improved
upon over time.
Conclusions: We conclude that this direction of research is worthy of further exploration.
AB - Background: Scientific research in bio-informatics is often data-driven and supported by biolog-
ical databases. In a growing number of research projects, researchers like to ask questions that
require the combination of information from more than one database. Most bio-informatics papers
do not detail the integration of different databases. As roughly 30% of all tasks in workflows are
data transformation tasks, database integration is an important issue.
Integrating multiple data sources can be difficult. As data sources are created, many design
decisions are made by their creators.
Methods: Our research is guided by two use cases: homologues, the representation and integration
of groupings; metabolomics integration, with a focus on the TCA cycle.
Results: We propose to approach the time consuming problem of integrating multiple biological
databases through the principles of ‘pay-as-you-go’ and ‘good-is-good-enough’. By assisting the
user in defining a knowledge base of data mapping rules, trust information and other evidence we
allow the user to focus on the work, and put in as little effort as is necessary for the integration.
Through user feedback on query results and trust assessments, the integration can be improved
upon over time.
Conclusions: We conclude that this direction of research is worthy of further exploration.
KW - EWI-22914
KW - METIS-296228
KW - IR-84256
M3 - Conference contribution
SN - not assigned
SP - 53
BT - BeNeLux Bioinformatics Conference 2012
PB - Centre for Molecular and Biomolecular Informatics
CY - Nijmegen
T2 - BeNeLux Bioinformatics Conference 2012
Y2 - 10 December 2012 through 11 December 2012
ER -