首页    期刊浏览 2024年09月12日 星期四
登录注册

文章基本信息

  • 标题:Capturing Data Provenance from Statistical Software
  • 本地全文:下载
  • 作者:George Charles Alter ; Jack Gager ; Pascal Heus
  • 期刊名称:International Journal of Digital Curation
  • 印刷版ISSN:1746-8256
  • 出版年度:2022
  • 卷号:16
  • 期号:1
  • DOI:10.2218/ijdc.v16i1.763
  • 语种:English
  • 出版社:University of Edinburgh
  • 摘要:We have created tools that automate one of the most burdensome aspects of documenting the provenance of research data: describing data transformations performed by statistical software.  Researchers in many fields use statistical software (SPSS, Stata, SAS, R, Python) for data transformation and data management as well as analysis.  The C2Metadata ("Continuous Capture of Metadata for Statistical Data") Project creates a metadata workflow paralleling the data management process by deriving provenance information from scripts used to manage and transform data.  C2Metadata differs from most previous data provenance initiatives by documenting transformations at the variable level rather than describing a sequence of opaque programs.  Command scripts for statistical software are translated into an independent Structured Data Transformation Language (SDTL), which serves as an intermediate language for describing data transformations.   SDTL can be used to add variable-level provenance to data catalogues and codebooks and to create "variable lineages" for auditing software operations.   Better data documentation makes research more transparent and expands the discovery and re-use of research data.
国家哲学社会科学文献中心版权所有