Why good data analysts need to be critical synthesists: Determining the role of semantics in data analysis

Simon Scheider, Frank Ostermann, Benjamin Adams

Research output: Contribution to journalArticleAcademicpeer-review

14 Citations (Scopus)

Abstract

In this article, we critically examine the role of semantic technology in data driven analysis. We explain why learning from data is more than just analyzing data, including also a number of essential synthetic parts that suggest a revision of George Box’s model of data analysis in statistics. We review arguments from statistical learning under uncertainty, workflow reproducibility, as well as from philosophy of science, and propose an alternative, synthetic learning model that takes into account semantic conflicts, observation, biased model and data selection, as well as interpretation into background knowledge. The model highlights and clarifies the different roles that semantic technology may have in fostering reproduction and reuse of data analysis across communities of practice under the conditions of informational uncertainty. We also investigate the role of semantic technology in current analysis and workflow tools, compare it with the requirements of our model, and conclude with a roadmap of 8 challenging research problems which currently seem largely unaddressed.
Original languageEnglish
Pages (from-to)11-22
Number of pages12
JournalFuture generation computer systems
Volume72
DOIs
Publication statusPublished - 2017

Keywords

  • METIS-322130
  • ITC-ISI-JOURNAL-ARTICLE

Fingerprint Dive into the research topics of 'Why good data analysts need to be critical synthesists: Determining the role of semantics in data analysis'. Together they form a unique fingerprint.

  • Cite this