The choice of pedagogy in statistics should take advantage of the quantitative capabilities and scientific background of the students. In this article, we propose a model for a statistics course that assumes student competency in calculus and a broadening knowledge in biology. We illustrate our methods and practices through examples from the curriculum.
When considering quantitative training for the next generation of scholars, the most persistently requested advanced skills expressed by the research-active life scientists at the University of Arizona are the use of statistics and comfort with calculus. Consequently, one centerpiece for the curricular activities at the University of Arizona is the development of three core mathematics courses–a two-semester-long sequence that integrates calculus and differential equations and, based upon that knowledge, a one-semester-long course in statistics.
Benefiting from changes in approach in the calculus and differential equations course, students enrolled in the statistics course are acquainted and have some facility with open-ended questions. In addition, because these students are typically juniors and seniors, they bring a much broader knowledge base in the life sciences than they had when they entered the calculus classroom for the first time.
Most life science students, presumed to be proficient in college algebra, are taught a variety of procedures and standard tests under a well-developed pedagogy. This approach is sufficiently refined so that students have a good intuitive understanding of the underlying principles presented in the course. However, if the needs presented by the science fall outside the battery of methods presented in the standard curriculum, then students are typically at a loss to adjust the procedures to accommodate the additional demand.
On the other hand, mathematics students frequently have a course in the theory of statistics as a part of their undergraduate program of study. In this case, the standard curriculum repeatedly finds itself close to the very practically minded subject that statistics is. However, the demands of the syllabus provide very little time to explore these applications with any sustained attention.
Our goal for life science students at the University of Arizona is to find a middle ground. In their overview “Undergraduate Statistics Education and the National Science Foundation,”
Despite the fact that calculus is a routine tool in the development of statistics, the benefits to students who have learned calculus are very rarely used in the statistics curriculum for undergraduate biology students. Our objective is to meet this need with a one-semester course in statistics that moves forward in recognition of the coherent body of knowledge provided by statistical theory having an eye consistently on the application of the subject. Even though such a course may not be able to achieve the same degree of completeness now presented by the two more standard courses described above, the question is whether it leaves the student capable of understanding what statistical thinking is, understanding how to integrate this with scientific procedures and quantitative modeling, how to ask statistics experts productive questions, and how to implement their ideas using statistical software and other computational tools. For a thoughtful essay on these issues, see
The efforts described above, despite having similar goals, arrive at very different approaches. In this article, we shall introduce the course at the University of Arizona with an annotated syllabus and through classroom examples, take home assignments, and end-of-the-semester projects.
The four parts of the course—organizing and collecting data, an introduction to probability, estimation procedures, and hypothesis testing—are the standard building blocks of many statistics courses. This section highlights some of the features of a calculus-based course.
Much of this is standard and essential—organizing categorical and quantitative data, appropriately displayed as bar charts, histograms, boxplots, time plots, and scatterplots, and summarized using medians, quartiles, means, weighted means, trimmed means, standard deviations, and regression lines. We use this as an opportunity to introduce students to the statistical software package R (
Collecting data under a good design is introduced early in the course, and discussion of the underlying principles of experimental design is an abiding issue throughout this course. With each new mathematical or statistical concept comes an enhanced understanding of what an experiment might uncover through a more sophisticated design than what was previously thought possible. The students are given readings on design of experiment and examples using R to create simple and stratified random samples. The recommended lecture is to tell a story that illustrates how these issues appear in personal research activities.
Probability theory is the analysis of random phenomena. It is built on the axioms of probability and is explored, for example, through the introduction of random variables. The goal of probability theory is to uncover properties arising from the phenomena under study. Statistics is devoted to the analysis of data. The goal of statistical theory is to articulate as well as possible what model of random phenomena underlies the production of the data. The focus of this section of the course is to develop those probabilistic ideas that relate most directly to the needs of statistics.
Thus, we must study the axioms of probability to the extent that the students understand conditional probability and independence. Conditional probability is necessary to develop Bayes formula, which we will later use to give a taste of the Bayesian approach to statistics. Independence will be needed to describe the likelihood function in the case of an experimental design that is based on independent observations. Densities for continuous random variables and mass function for discrete random variables are necessary to write these likelihood functions explicitly. Expectation will be used to standardize a sample sum or sample mean and to perform method of moments estimates.
Random variables are developed for a variety of reasons. Some, like the Poisson random variable or the gamma random variable, arise from considerations based on Bernoulli trials or exponential waiting. The hypergeometric random variable helps us understand the difference between sampling with and without replacement. The
The flavor of the course returns to becoming more authentically statistical with the law of large numbers and the central limit theorem. These are largely developed using simulation explorations and first applied to simple Monte Carlo techniques and importance sampling to evaluate integrals. One cautionary tale is an example of the failure of these simulation techniques when applied without careful analysis. If one uses, for example, Cauchy random variables in the evaluation of some quantity, then the simulated sample means can appear to be converging only to experience an abrupt and unpredictable jump. The lack of convergence of an improper integral reveals the difficulty.
The central object of study is, of course, the central limit theorem. It is developed both in terms of sample sums and sample means and used in relatively standard ways to estimate probabilities. However, in this course, we can introduce the delta method, which adds ideas associated to the central limit theorem to the context of propagation of error.
In the simplest possible terms, the goal of estimation theory is to answer the question:
The point estimation techniques are followed by interval estimation and, notably, by confidence intervals. This brings us to the familiar one- and two-sample
For hypothesis testing, we begin with the central issues—null and alternative hypotheses, type I and type II errors, test statistics and critical regions, significance, and power. We begin with the ideas of likelihood ratio tests as best tests for a simple hypothesis. This is motivated by a game. Extensions of this result, known as the Neyman Pearson lemma, form the basis for the
The desire of a powerful test is articulated in a variety of ways. In engineering terms, power is called sensitivity. We illustrate this with a radon detector. An insensitive instrument is a risky purchase. This can be either because the instrument is substandard in the detection of fluctuations or poor in the statistical test that results in an algorithm to announce a change in radon level. An insensitive detector has the undesirable property of not sounding its alarm when the radon level has indeed risen.
The course ends by looking at the logic of hypotheses testing and the results of different likelihood ratio analyses applied to a variety of experimental designs. The delta method allows us to extend the resulting test statistics to multivariate nonlinear transformations of the data.
In this section, we describe three examples from the University of Arizona course. These examples are presented to highlight how a calculus-based course differs from an algebra-based course. These abbreviated descriptions cannot bring the same sense of background preparation or the breadth of issues under consideration that a student experiences in the classroom. Consequently, more detailed notes are available from the author upon request.
For the first two examples, we begin with the presentation seen in a typical algebra-based statistics course and then provide extensions of these ideas that are possible for those students who have good calculus skills. The third example describes a strategy to introduce the concept of likelihood. In subsequent classes, the students will use this as motivation for the rationale for the optimization problems in the development of likelihood ratio tests.
Given observations ( the size of the residuals depends on the sign of the residuals depends on
The goal here is to have the ability to extend the use of regression beyond the formulas and qualitative reasoning. To address the first case above, we can modify the least squares criterion and solve a weighted least squares regression with weight function
Here, we focus on the second case in which the scatterplots display a curved relationship. With the tools of college algebra, this relationship can be transformed to be suitable for linear regression by guessing the relationship and applying the inverse transformation to the response variable. This is seen most frequently in scatterplots of a quantity over time exhibiting exponential growth or decay. Thus, we take the logarithm of the response variable and proceed.
This strategy is sometimes too superficial to be successful. For students who have been acquainted with differential equations, we can consider a common chemical reaction,
Michaelis–Menten kinetics. Measurements of production rate
However, this classical approach to estimation is no longer considered satisfactory (
One of the aspects of any new course is the understandable uncertainty by the students that the approaches presented in class have applicability to questions of their own concern. Thus, in the problem sets, we ask the students to use the ideas presented in lecture and make substantial use of them in another biological context. In this case, the ideas on nonlinear transformations are further explored by the students in the following question addressed by
Due to selection, neutral sites near genes are hitchhiked along with the gene to higher levels in the population until a recombination event separates the neutral site from the selected gene. This reduces the diversity of the genome near genes. The nucleotide diversity π versus recombination rate ρ is given for 17 gene regions in
Double reciprocal plot for genetic hitchhiking. 1/π versus 1/ρ. The regression line
George
Unfortunately, the best proofs of the central limit theorem rely on indirect and sophisticated methods. Moreover, these proofs yield very little insight into the emergence, with an increasing number of observations, of the bell curve for the distribution of
Most elementary courses dedicate a week to the understanding of the central limit theorem. This development might culminate with a graph like that seen in
Displaying the central limit theorem graphically. Density of the standardized version of the sum of
The extension we can have in a calculus-based course is motivated by interest in nonlinear transformations of the measured quantities. Some physics, chemistry, and engineering students have seen a bit of this in the study of propagation of error (see, e.g.,
After investigating some one-dimensional examples like those seen in
Illustrating the delta method. Here the mean μ =
A critical value for
The central limit theorem can directly determine appropriate normal distributions to approximate F̄, the sample mean of the number of female fledglings per successful nest, p̂, the sample proportion of surviving nests, and N̄, the sample mean number of nests built per female per year. If our estimate B̂ of the fecundity is the product of these three numbers, then what is a good approximation for the distribution of this statistic?
Propagation of error analysis suggests that we find a linear approximation to
The delta method extends the propagation of error analysis by noting that, by the central limit theorem, the random quantity given by the linear approximation can be approximated by a normal distribution. Returning to the question of estimating fecundity, the delta method tells us that B̂ can be approximated by a normal distribution. The variance can be determined from σF2, the variance in the fecundity measurement and σN2, the variance in the number of nests built per adult female per year. By combining this information, we derive in the class an expression that highlights in the contribution to the variance in the estimate of fecundity originating from each of the three measurements:
This application of the calculus greatly extends the applicability of the central limit theorem in the practical application of statistics. This idea is explored by the students who apply the delta method to describe the estimator for the focal length
Simulating a sampling distribution for the estimate of focal length. In this example, the distance from a convex lens to an object
The appearances of the delta method do not end here. We shall use this technique to determine the variance for a method of moments estimator. Later, we can use these ideas to construct test statistics for hypothesis tests beyond those generally encountered in more elementary courses.
The study of statistical hypotheses begins with a shot of jargon that will take a student some time to absorb. Understanding the terminology is important because it codifies the relationship between hypothesis testing and scientific advancement. Under a simple hypothesis, we write the test as:
Reject the hypothesis. Rejecting the hypothesis when it is true is called a type I error or a false positive. Its probability α is called the size of the test or the significance level. Fail to reject the hypothesis. Failing to reject the hypothesis when it is false is called a type II error or a false negative. If the false negative probability is β, then the power of the test is 1 − β. > x<-c(−11:11) > L0<-c(0,0:10,9:0,0) > L1<-sample(L0,length(L0)) > data.frame(x,L0,L1)
The rejection of the hypothesis is based on whether or not the data
The language of hypothesis testing.
Thus
The goal of this game is to pick values
Being ahead by a score of 23–0 can be translated into a best critical region in the following way. If we take as our critical region
Understanding the next choice is crucial. Candidates are
Results for Neyman-Peason game.
From this exercise we see how the likelihood ratio test is the choice for a most powerful test. This is more carefully described in the Neyman-Pearson lemma [See
Let
Using R, we can complete the table for > o<-order(L1/L0) > sumL0<-cumsum(L0[o]) > sumL1<-cumsum(L1[o]) > alpha<-1-sumL0/100 > beta<-sumL1/100 > data.frame(x[o],L0[o],L1[o], + sumL0,sumL1,alpha,1-beta)
Completing the curve, known as the receiver operator characteristic (ROC), is shown in
Receiver operator characteristic. The graph of
Now the students are prepared to understand the reasoning behind the use of critical values for the
These examples have been presented here in such a way as to emphasize transitions from the use of algebra to the use of calculus in the learning of concepts in statistics. However, students do not notice such a shift in the sense that they do not see calculus as a special tool. Limits need to be evaluated, rates change, areas under curves need to be determined, functions need to be maximized, and their concavity needs to be assessed. Moreover, calculus is more than a powerful computational tool: it provides a way of thinking that enhances the students' view of what is possible. I asked a student how much is calculus used in the course. The answer she gave, “So much we don't think about it,“ is a goal for this course.
For many years, quantitative subjects as diverse as physics and economics have developed distinct pedagogies for an algebra-based and for a calculus-based course. For example, Newton's second law applied to projectile motion or to the motion of springs gives, for the calculus student, simple differential equations from which the basic algebraic or trigonometric relationships for the variables position, velocity, acceleration, and time can be derived. For the algebra student, these relationships have to be taken on faith. In a similar way, students can receive a more comprehensive and elegant understanding of the fundamental principles of statistical science if they are equipped with a working knowledge of calculus.
One important feature of the course is the end-of-the-semester project. Typically for this assignment, students work in pairs. All of the suggested projects are based on research activities that have taken place at the University of Arizona, and all of the projects require statistical analyses beyond the methods presented in the course. This gives the students the opportunity to work as a team on a project and to work with one of the scientists (typically a graduate student or postdoctoral fellow) who was involved in the original research. This way, the students obtain a first-hand account of the fundamental questions that are meant to be addressed by the research and learn how to use the ideas from their statistics course and apply them to new situations. They certainly see the nature of open-ended questions that are a part of a research scientist's daily life.
We describe two projects to give a sense of the breadth of the choices from the point of view of both biology and statistics.
Human populations and the languages that they speak change over time. The movement of people and the innovations they make in their languages are difficult to observe and quantify over short periods and impossible to witness over long periods. Consequently, researchers have been forced to undertake indirect approaches to infer associations between human languages and human genes. Many well-known studies focus their questions on the movement of people on continental scales. These studies led
Using the Indonesian island of Sumba as a test case,
To investigate the paternal histories of the Sumbanese,
The linguistic data consist of 29 200-word Swadesh lists from sites well distributed throughout the island. Swadesh lists are built from words ascribed to meanings that are basic to everyday life. Using the methodology of comparative linguistics, some words from different language lists can be traced to a common ancestral word. These techniques have been used previously to construct a Proto-Austronesian (PAn) language. In addition, a phylogenetic tree built from these 29 word lists branches to form five major language subgroups and leads us to the conclusion that the present day Sumbanese all speak a language derived from a single common ancestral Proto-Sumbanese. From this we can count the number of words on the Swadesh word lists that are derived from the PAn language. This information is summarized in
Phylogenetic and geographic distribution of languages and Y chromosome haplogroups.
The Mantel test (
In this example, we have
To see whether the value of
Mantel test results for correlation.
We next test the null assumption of no correlation between the fraction of O haplogroup men and retained Swadesh list PAn cognates among the eight villages. The alternative is that these two quantities are positively correlated. A naive strategy to analyze the correlation in
Scatterplot of PAn cognates versus the percentage of sample from haplogroup O. The correlation is 0.627 (
An exact or even an approximate computation of the distribution of correlation under the null hypothesis is difficult. Consequently a bootstrap analysis (
For each of the resulting bootstrap samples, the students can compute the correlation. Repeat this procedure many times to create the bootstrap distribution of correlation under the null hypothesis. The bootstrapped
One of the most widely used relations in chemistry is the Arrhenius relation,
Single molecule experiments. Set-up for kinesin attached to a bead, translocating along a microtubule subject to optical tweezers pulling on bead. (
Taking displacement to be our reaction coordinate, the free energy Δ
At the level of a single interaction, then for a simple reaction,
The dwell time is based on thermal fluctuations and thus possesses the memorylessness property. In other words, the dwell times are random and follow an exponential distribution (see The Arrenhius relation holds at the level of a single molecule, i.e., the mean dwell time equals
Previously published methods used seat of the pants bootstrapping techniques. The graduate student who brought this problem to my attention performed some numerical experimentation, and it appears that the estimators do not even converge as the number of observations increases. However, students can apply the concepts in the course to design most powerful tests and with it asymptotically narrowest possible confidence intervals.
If our data are from an experiment with independently measured forces
This experiment has not yet been performed (Kalafut, Liang, Watkins, and Visscher, unpublished results), so we ask the students to simulate data and find the maximum likelihood estimates. This gives us the opportunity to discuss the value of generating data via simulation and test the inference methods on these data before moving to the actual data. To keep the situation simple, we take δ = 1 and τ0 = 1. We also take the forces uniformly distributed between 0 and
Force versus dwell time for the simulated data. The Arrhenius relationship is not easy to visualize in the scatterplot of 1000 observations.
To simulate the data with > F<-runif(1000); t<-rep(0,1000) > for(k in 1:1000){t[k]<-rexp(1,exp(F[k]))} > plot(F,t)
If we solve numerically for the estimates in this simulation, we obtain
Maximum likelihood estimation. The two curves are the two expressions for
The covariance matrix for
The remaining three projects are drawn from questions on bacterial growth and division, on otoacoustic emissions in the human ear, and on flour beetle population dynamics. The expectation is that if the student has a genuine interest in the life sciences, then at least one of these projects will be enticing.
At the University of Arizona, we are seeing a transformation in the view that mathematics plays in the education of life scientist students. To see this evolution reflected in the statistics class, we saw six students choose to add a mathematics minor during the first time the course was given. For the second and third time, the students entered the class as mathematics minors. The course, offered for a fourth time this fall, is full and the capacity is being increased to accommodate demand.
The head of our math center is now writing to students who have completed the first two years of mathematics course work and is encouraging them to take the calculus-based statistics course. In addition, more than a fourth of the students (26 of 98) in the Undergraduate Biology Research Program are either mathematics majors or minors. Several researchers who have had statistics students in their lab are recommending it to other students.
From the course evaluations, the students especially liked “the applicability of statistics to multiple areas of interest,” that “the homework sets were challenging and engaging,” “relating course work to software used in research,” having “taken stats classes before I feel like I actually learned some of the basics of stats in this one,” “learn(ing) about the real world,” “forc(ing) me to learn a subject I really didn't enjoy (I like stats now),” that “the study examples were particularly effective as motivation to learn the material,” “using R and relating everything to the real world,” and that “the homework was nicely challenging.” Thus, we can see that students are favorably impressed by the applicability of the course and the ability of theoretical considerations to lead to strategies for addressing practical problems. After some gentle reminding, they recognize that the material in some of the more mathy courses in their past were essential in their ability to address life science questions with their present level of understanding. Students are showing more facility with software as they continue to use it and having a general sense that mathematical and statistical issues are aspects of the challenges that come with exciting research.
The students, especially those who are working in a laboratory, were invited to select their own topic of interest. About a third of the students picked this option. Yeast genetics, earthquake prediction, rat behavior, fruit fly behavior, strength of material, educational value of certain curricula and evaluation methods, land subsidence, malaria prevalence, and the effects of monetary policy were among the chosen topics. Most of these student-generated projects resulted in a presentation in their research group's lab meeting.
To get a sense of the source and motivation of this transformation, we are now systematically gathering information on changes in student and educator attitude as a consequence of a variety of experiences, including this course. Analysis of these data will bring further insights into the type of attitudes that students who choose this course have and how their attitude is impacted by the experiences of the course.
The
As the course described in this article settles on an approach and a curriculum, we can now see that the uses of calculus in these contexts turn out to be neither particularly difficult nor novel for a student comfortable with the subject. In addition, after a couple of frustrating weeks with syntax, students continue acquiring skills in the use of statistical software. What is difficult for the students is the increased maturity necessary to apprehend the breadth of applicability that statistics brings to any carefully conceived data-rich exploration. In the end, having the tools of algebra and calculus, with probability as a frame of reference and ready access to computational software, students can kindle excitement in their new discoveries in the life science and bring themselves to a level of understanding that is not accessible absent the combined use of these quantitative tools.
Despite the aspiration inherent in the title
I thank the Howard Hughes Medical Institute (HHMI) Biomath Committee at the University of Arizona for many thought-provoking discussions. Student work and their feedback were essential in setting and refining the pedagogical approaches in the course. A special note of thanks goes to Christopher Bergevin, who was co-instructor the first time the course was taught, and to Carol Bender, the Director of the Undergraduate Biology Research Program, whose efforts have provided many undergraduate students with the opportunity to experience an authentic research experience. The activities described in this article were supported in part by a grant to the University of Arizona from the HHMI (52005889).