EAS 501.071 - Environmental Data Science
Data science is a rapidly changing area of environmental scholarship. Environmental data are growing in size, diversity, and availability at an unprecedented rate. This data deluge creates new opportunities for discovery, but also new challenges: finding relevant data, understanding data structures, cleaning and transforming data, documenting analytical choices, communicating results, and producing workflows that others can inspect, reproduce, and extend.
In this course, students will learn the fundamental practices and principles of environmental data science. The course primarily uses the R programming language, along with tools such as Quarto, Git, GitHub, and SQL. While these tools provide the technical foundation for the class, the course emphasizes broader data science principles that transfer across software environments, including reproducibility, visual encoding, literate programming, provenance, tidy data, relational data models, abstraction, automation, temporal structure, and spatial data models.
The class is organized around hands-on modules supported by environmental applications such as climate data exploration, grassland experiments, CO₂ records, forest inventory data, carbon accounting, ecological text mining, phenology, and marine habitat suitability. Each module connects a computational topic to a broader principle and an applied environmental example. Students will use these skills throughout the semester to design an individual project related to environment and sustainability, identify or assemble data, perform analyses, communicate results, receive feedback, and complete a final paper.