Books and Articles by David Semmelroth

David Semmelroth has two decades of experience translating customer data into actionable insights across the financial services, travel, and entertainment industries. David has consulted for Cedar Fair, Wachovia, National City, and TD Bank.

Articles & Books From David Semmelroth

Cheat Sheet / Updated 03-10-2022
Summary statistical measures represent the key properties of a sample or population as a single numerical value. This has the advantage of providing important information in a very compact form. It also simplifies comparing multiple samples or populations. Summary statistical measures can be divided into three types: measures of central tendency, measures of central dispersion, and measures of association.
Article / Updated 03-26-2016
There are several Exploratory Data Analysis (EDA) techniques you can use to test assumptions about a dataset. These include run sequence plot, lag plot, histogram, and normal probability plot. Run sequence plot Many statistical techniques are based on the assumption that the data being analyzed has the following properties: Independent variables Variables drawn from a common probability distribution Variables with common parameters (for example, mean and standard deviation) A run sequence plot tests whether the data conforms to these assumptions.
Article / Updated 03-26-2016
Measures of central tendency show the center of a data set. Three of the most commonly used measures of central tendency are the mean, median, and mode. Mean Mean is another word for average. Here is the formula for computing the mean of a sample: With this formula, you compute the sample mean by simply adding up all the elements in the sample and then dividing by the number of elements in the sample.
Article / Updated 03-26-2016
Regression analysis is used to estimate the strength and direction of the relationship between variables that are linearly related to each other. Two variables X and Y are said to be linearly related if the relationship between them can be written in the form Y = mX + b where m is the slope, or the change in Y due to a given change in X b is the intercept, or the value of Y when X = 0 As an example of regression analysis, suppose a corporation wants to determine whether its advertising expenditures are actually increasing profits, and if so, by how much.
Article / Updated 03-26-2016
One technique you can use to identify the distribution a dataset follows is the QQ-plot (QQ stands for quantile-quantile). You can use the QQ-plot to compare a dataset to a large number of different probability distributions. Often, data is compared to the normal distribution because many statistical tests assume normally distributed data.
Article / Updated 03-26-2016
Prior to performing any type of statistical analysis, understanding the nature of the data being analyzed is essential. You can use EDA to identify the properties of a dataset to determine the most appropriate statistical methods to apply to the data. You can investigate several types of properties with EDA techniques, including the following: The center of the data The spread among the members of the data The skewness of the data The probability distribution the data follows The correlation among the elements in the dataset Whether or not the parameters of the data are constant over time The presence of outliers in the data Another key question EDA answers is "Does the data conform to our assumptions?
Article / Updated 03-26-2016
A quantile-quantile plot (also known as a QQ-plot) is another way you can determine whether a dataset matches a specified probability distribution. QQ-plots are often used to determine whether a dataset is normally distributed. Graphically, the QQ-plot is very different from a histogram. As the name suggests, the horizontal and vertical axes of a QQ-plot are used to show quantiles.
Article / Updated 03-26-2016
A time series is a set of observations of a single variable collected over time. With time series analysis, you can use the statistical properties of a time series to predict the future values of a variable. There are many types of models that may be developed to explain and predict the behavior of a time series.
Article / Updated 03-26-2016
In data analysis, the relationship between the mean and the median can be used to determine if a distribution is skewed. The histogram shows that most of the returns are close to the mean, which is 0.000632 (0.0632 percent). The median is −0.0001179. Histogram shows most returns close to the mean. Here's how to determine whether the distribution is skewed: In this case, the distribution of returns to ExxonMobil stock is positively skewed.
Article / Updated 03-26-2016
Data is stored in different ways in different systems. So it's no surprise that when collecting and consolidating data from various sources, it's possible that duplicates pop up. In particular, what makes an individual record unique is different for different systems. An investment account summary is attached to an account number.