Skip to contents

HistData collects small data sets that are interesting or important in the history of statistics and data visualization. Contributions are welcome: new data sets, corrections to existing ones, better documentation, and examples that re-create or re-think a historical analysis or graphic.

The simplest way to propose something is to open an issue describing the data and where it comes from, before doing much work. Pull requests are welcome once we have agreed it fits.

What makes a good HistData data set

  • It has a place in the history of statistics or data visualization: the data behind a landmark analysis, a famous graphic, or a first use of a method.
  • It is small enough to ship in a CRAN package and to read in a help page.
  • Its source can be identified and cited, so that someone else could check the numbers.
  • It offers something to do: an analysis to reproduce, a graph to re-create, or a question the original author left open.

Please contribute only material that you have the right to share and that can be redistributed under the package’s license (GPL).

Data. Most data here was transcribed from historical publications that are in the public domain, and that is the safest kind of source. If your data comes from somewhere else, check before contributing:

  • A modern book, article or supplement: the numbers themselves are generally not protected in the way the text and layout are, but a modern compilation, database or cleaned-up version may carry its own rights or license terms. Say what the terms are.
  • Another R package, a GitHub repository or a website: check its license, and credit it. “It was on the web” is not permission.
  • Someone else’s transcription, digitization or correction of a historical source: this is their work. Ask them, and say in the documentation that you did. If permission is pending, say so and keep the file out of the repository until it is given.

Images. Do not add scans, photographs or reproductions of graphics unless they are in the public domain or you have permission to redistribute them. A library’s scan of an old book may come with its own terms of use, so check and record them. Where possible, link to the image at its source rather than copying it into the package. Examples should draw their own graphs from the data.

Text. Write documentation in your own words. Short quotations with a citation are fine; do not paste long passages from copyrighted books, articles or translations.

If you are unsure whether something can be included, open an issue and ask. Nothing here is legal advice; it is the standard this project tries to hold itself to.

Provenance

A data set’s documentation should trace a clear path from the original source to the version included here. For each contribution, please record:

  • The original source, with a full citation, and a link to a scan or digital copy where one exists.
  • How the data got from there to here: transcribed by hand, read from a scan, digitized from a graph, taken from an intermediate source. If it came through an intermediate source, name it.
  • What you changed: corrections, recoding, reshaping, computed variables.
  • What is uncertain: illegible values, misprints in the original, figures that do not add up, anything you could not verify. Say so plainly; it is better than implying more confidence than the data warrants.

Keep the raw transcription and the script that turns it into the R data set in data-raw/, so the work can be checked and repeated.

Corrections

If you find an error in an existing data set, please open an issue that gives the value in the package, the value you believe is correct, and the source you checked it against (with a page or table number). Differences between the package and a later secondary source are worth reporting too, even when the package turns out to be right.

Preparing a contribution

  • Data sets are saved as .RData files in data/, one object per file.
  • Documentation is written with roxygen2, using markdown, in a file in R/ named for the data set. Include @format, @source, @references, one or more @concept tags for the statistical or graphical ideas the data illustrates, and @examples that run.
  • Do not edit files in man/ by hand; they are generated by devtools::document().
  • Examples should run quickly and use base R or packages already listed in Suggests.
  • Add a line to NEWS.md.

Credit

Contributors are acknowledged in the documentation of the data sets they worked on, in NEWS.md, and, for substantial contributions, in the package DESCRIPTION.

Conduct

Please be courteous and constructive in issues and pull requests. This project also has a Code of Conduct.