CNLS 503 Final Exam | Complete Solutions, Latest Version True of False: Statistics is used to prove claims. False What are Variables? Variables are measurable characteristics that can vary in value. My answer: Variabl
...
CNLS 503 Final Exam | Complete Solutions, Latest Version True of False: Statistics is used to prove claims. False What are Variables? Variables are measurable characteristics that can vary in value. My answer: Variables are different types of data in an experiment (used to classify the type of data being collected) What type of scale is used when labeling or using categories? This would be qualitative data and the SCALE would be nominal! Ex: blood type, college majors, room numbers True or false: it is generally feasible to collect data from every member of a population FALSE! unless the target of a particular research study is comprised of a small population, it is usually not practical or sensible to attempt to reach every member within a population Which of the following are examples of ordinal variables? Select all that apply. - Shirt sizes - college major - places in a marathon (1st, 2nd, 3rd) - Temperature Explanation: ordinal is used for qualitative variables and includes ranking in size or measure (does not indicate how much the data differ) When a variable can contain negative values, what type of scale is most appropriate? Why? An interval scale because variables containing negative values do not have an absolute zero. Note: this is quantitative data When is it most appropriate to use a ratio scale? A ratio scale is used when a variable has an absolute zero (Ex: height, weight, exam scores, GPA) Note: this is quantitative data True or false: statistics are used to communicate results but can be used to mislead individuals TRUE! Provide an example of quantitative variable and an example of a qualitative variable Quantitative Ex: test scores, reaction time, error number, vital signs, temperature Qualitative Ex: sex, college major, eye color, blood type, shirt size How are sample and population related? A population includes every member within a particular group, whereas a sample is a smaller subset of the given population (used to represent the whole population) True or false: discrete variables are counted as whole numbers TRUE! discrete variables are quantitative variables that can be organized into separate categories or counted using whole numbers. Ex: number of patients in a hospital, pets in a home _________ statistics involves analyses that provide a way to summarize and describe data, whereas _________ statistics are performed for researchers to make inferences and generalizations about populations based on sample data. descriptive; inferential Explanation: descriptive statistics involve calculations like mean, median, mode, standard deviation, variance, and range. Inferential statistics involve calculations such as those use for hypothesis testing (z-scorem, t-test, ANOVA, correctional analyses). Which of the following is NOT involved in the process of sampling? - identify goals - collect data - make conclusions - obtain resources Explanation: the steps are identify goals, gather sample, make inferences/conclusions A movie theater is interested in how movie-goers would rate the current movies being shown. They survey a group of 100 movie-goers and ask each to rate the movie they saw on a scale of 1-5 (1 = terrible, 5 = excellent). The surveyors then calculate the average rating of each movie. This is an example of what kind of statistic? - descriptive statistics - sample statistics - parametric statistics - inferential statistics descriptive statistic because average rating of movie is a calculation used to summarize / describe Which of the following is a discrete variable? Select all that apply. - Number of pets per household - distance - number of patients at a hospital - exam scores because discrete variables can ONLY be counted as whole numbers! Name a continuous variable. Why is it considered continuous? Ex: weight, temperature, time, length, and speed. continuous variables are variables that can be broken down into smaller, fractional components (decimals, fractions, etc). It does NOT have to be a whole number Which of the following fields do NOT use statistics? - politics - sports - entertainment - medicine - none of these fields use statistics - all of these fields use statistics Which of the following statements is TRUE regarding populations? - it is expected that every member of a population will agree to participate in a study. - studying every member of a population is how most behavioral and clinical studies are conducted - it is only practical to reach all members of small populations - only one method is used to gather samples from populations A team of clinical researchers is studying stress levels of a sample of 15 paramedics undergoing a new training program at a local hospital. The researchers stage two hypothetical medical emergencies for the paramedics: one low stakes and one high stakes. Then they measure their cortisol levels. The scenarios are administered at the same time every day, but the location changes based on the scenario being presented. In this experiment - identify the independent, dependent, control, and extraneous variables. Independent: emergency situations (scenario) Dependent: cortisol levels Control: time of day Extraneous variables: location Note: Extraneous variables are any variable not being investigated that has the potential to affect the outcome of a research study. Control variable is an experimental element which is constant (controlled) and unchanged throughout! Why is the field of statistics important? it allows data to be described and communicated succinctly and concisely. Statistics allow inferences to be drawn about data, particularly when it is not feasible to collect information from all members of a certain group. Statistics equip us with the necessary tools needed to critically evaluate information. My answer: statistics gives people a guide in a world where there is so much information - news, advertisements, products, etc. This gives people a way of sorting through data, organizing numbers, and making everyday decisions. when the frequency for every value in a dataset is presented, a __________ is formed: - frequency diffusion - frequency distribution - frequency dispersion - frequency dispersal a range of scores can be grouped in sets called ____________. bins groups categories chunks Explanation: when there is a wide range of values a grouped frequency distribution table can be used. Bins are used! calculate the range for the following scores: 45, 64, 32, 58, 43 Range = Xmax - Xmin range = 64 -32 =32 Which of the following is a measure of variability? - mean - median - range - mode what is central tendency? one way to describe data; it is an average or middle value that describes the center of a distribution. Mean, median, and mode are three measures of central tendency True or false: outliers are always the result of errors in data collection FALSE! outliers may be legitimate values that should be included in the results True or false: frequency tables are used to display both qualitative and quantitative data. TRUE! displays the number of times a certain value appears within the dataset (frequency). One of the easiest ways to summarize a set of score or values. identify the mode in the following set of scores: 438, 421, 423, 432, 440, 435, 441 there is no mode because every value appears the same amount of times (once) Why are pie charts not used frequently be researchers and statisticians? the effectiveness of pie charts is limited to datasets with only a few categories My answer: they are best with a small number of relative frequencies - too many "slices" make it difficult to interpret the pie chart the sum of all relative frequencies in a dataset will always equal ____________. 1 or 100% Calculate the mean for the following set of scores: 89, 75, 91, 68, 72, 83, 94, 78 mean (x̄) = Σ (sum of) X / n x̄ = (89 + 75 + 91 + 68 + 72 + 83 + 94 + 78) / 8 = 81.25 true or false: the median is affected by outliers FALSE! Neither the mode nor median is affected by outliers. It IS possible for the mean to be affected. a neuropsychologist records the scores on the MMSE (Mini-Mental State Examination) for 30 patients. Every score appears once in the dataset. What would this distribution be known as? - multimodal distribution - uniform distribution - bimodal distribution - unimodal distribution Explanation: distribution can be described by shape and symmetry. The number of peaks is determined by # number of modes. Since there is no mode it is uniform distribution. determine the median for the following set of scores: 2121, 2115, 2117, 2120, 2118, 2122 First put in order from least to greatest (2115, 2117, 2118, 2120, 2121, 2122). Then find the middle number. If there is two take the mean of the two. (2118 + 2120) / 2 = 2119 the bars of __________ do not touch because they represent discrete values, whereas the bars of ________ do touch because they represent continuous values bar graph; histogram Explanation: each bar in a bar graph/bar chart represents an individual category. Aaron scores a mean of 85 (s = 8) on 4 of his cognitive psychology exams for fall semester. Kyle scores a mean of 83 (s = 4) on the cognitive psychology exams. Does Aaron or Kyle have the more consistent performance on the exams? Kyle because his standard deviation (s) was lower, which indicates less variability in his test scores compared to Aaron Which type of graph clearly depicts the shape of a distribution? FREQUENCY POLYGON because the increasing and decreasing sections of the data can be seen In a symmetrical distribution, what is the most commonly used measure of central tendency? MEAN Note: in a symmetrical distribution, the mean/median/mode are all the same value researchers are studying the average miles per gallon (mpg) of different cars based on size (compact, mid-size, and full-size) and make (Ford, Chevy, and Chrysler). Would a three dimensional graph be appropriate to use to display these data? Explain. Yes because there are three variables being measured (mpg, car size, and car make). So a three dimensional graph would be the best choice. three dimensional graph is best for "three dimensional data" seven food trucks are competing to see how many customers visit their trucks per day in a busy part of the city. The average number of customers that visit each food truck per day is 281. The sum of squares is found to be 187. calculate the variance for these data s² or σ² (variance) = Σ (X - x̄)^2 / (n - 1) and Σ (X - x̄)^2 = SS (sum of squares) variance = sum of (term in data set - sample mean) squared / (sample size - 1) sum of squares = 187 s² = 187 / (7-1) = 31.17 a gardener has 5 apple trees in her yard. Each tree has a different number of apples growing on it. The following list represents the number of apples on each tree: 22, 29, 31, 23, 27. Calculate the standard deviation of these values (x̄ = 26.4) Standard deviation (s) = √ (variance) √ Σ (X - x̄)^2 / (n - 1) and Σ (X - x̄)^2 = SS. So, Σ (19.36 + 6.76 + 21.16 + 11.56 + 0.36) = 59.2 √ (59.2 / 5-1) = √ 14.8 = 3.85 a professor posts the final exam grades for 50 of his psychology students. Most of the scores are clustered on the lower side of the distribution, with a few scores falling toward the right. What term would be best describe the distribution? what does this imply about how hard the test was? this would be considered a positively skewed (right-skewed) distribution, as most of the scores fall on the left side of the distribution and the skew or "tail" is over the right side. This would indicate that the test was very difficult, as most students scored low on the exam true or false: a statistic applies to a population, and a parameter applies to a sample FALSE! A statistic applies to a sample, whereas a parameter applies to a population Remember: s to s and p to p true or false: cluster sampling involves dividing a population into subgroups and then randomly selecting several groups for the study TRUE! Cluster sampling involves dividing a population (stratum) into "clusters" (subgroups) then randomly selecting several groups. Better than simple random sampling and stratified sampling when there is a very large population. What causes sampling bias? sample bias generally occurs when the researcher has a mistake in the data collection or measurement process it can be from biased sample choice (convenient sampling), selection bias, errors made while collecting the data (measurement bias), response bias, miscalculations, etc. this type of sampling involves randomly gathering data from subgroups in the population known as strata STRATIFIED SAMPLING! It's a type of random sampling used when the population can be divided into subgroups (called strata). Used when a researcher wants to compare outcomes for different subgroups w/in a population or to compare outcomes between subgroups. what is a statistical hypothesis? a claim made about a population parameter the area in the distribution of sample means where a low probability exists is called ____________ - critical region - hypothetical region - skewed region - significant region If a test statistic lies in the critical region then there is reason to reject the null hypothesis. generally, if a test statistic is ___________ than the critical value, then it is considered statistically significant. - less extreme - better - more extreme - more central "more extreme" which is the same as saying greater than or exceeds. regional managers administer a job satisfaction survey to their employees. Some of the employees are fearful that if they respond negatively to the questions, then they will be fired. So, they provide inaccurate responses. What type of bias is occurring here? - response - selection - measurement - proprietary this occurs when participants respond to surveys with inaccurate, untruthful, or exaggerated responses. as the sample size increases, does the sample mean more closely or less closely represent the population mean? MORE CLOSELY! This is summarized by the central limit theorem. It states that the mean of the distribution of sample means is equivalent to the population mean for large sample sizes ( n = 30 or more) & the distribution of sample means is an approximately normal distribution for large sample sizes. sample sizes with less than __________ members are considered small. 30 50 100 20 the ___________ involves performing numerous observations of a given situation and recording the number of times an event occurs RELATIVE FREQUENCY METHOD! Then a frequency distribution graph can be used to display the distribution of scores. a _____________ is a frequency distribution of each statistic from every possible sample of a given size from the population sampling distribution a bag of marbles contains 42 red balls, 38 green balls, and 23 yellow balls. Calculate the probability of selecting a green ball from the bag. probability (A) = number of outcomes in A / total number of possible outcomes. 38 / 103 = 0.37 what term refers to the average error expected between the sample mean (x̄) and the population mean (μ) in a sampling distribution? STANDARD ERROR ( 𝜎x̄ ). Also known as standard deviation. Researchers are studying the effects of a new cognitive-behavioral therapy (CBT) technique on anxiety. They recruit a sample of 275 patients with anxiety, and the patients participate in the therapy. The researchers then compare the mean anxiety levels of the sample to the population mean. What would the alternative hypothesis for this study state? Ha : there is a relationship between the CBT technique and anxiety levels the fresh fruit section in a grocery store is comprised of 30% citrus fruits, 40% berries, 15% apples, 10% melons, 5% bananas. You purchase 10 pieces of fruit, with 4 citrus fruits, 2 packs of strawberries, 2 watermelons, 1 apple, and 1 set of bananas. Is your purchase a representative or biased sample of the fresh fruit section? Explain your answer. The sample is comprised of 40% citrus fruits, 20% berries, 10% apples, 20% melons, 10% bananas. Therefore, this is a BIASED sample because it is not proportionate or representative of the original population. researchers use a random number generator to select a sample of 40 individuals. What type of sampling method is being used? SIMPLE RANDOM SAMPLING! This is when each individual in the population has an equal change of being selected for the sample. true or false: convenience sampling typically results in a biased sample TRUE! For a sample size of 100, find the standard error for a population with μ = 150 and σ = 11. σ (x̄) = σ / √ n standard error of population (sample) = standard deviation / √ sample size 11 / √ 100 = 1.10 in hypothesis testing, what result is required for the null hypothesis to be rejected? the test statistic falls in the critical region; this means the test statistic is compared to the critical value and it is greater than the critical value (this shows statistical significance) Note: test statistics can be used to find p-value which is then compared to the alpha level which of the following are principles of the central limit theorem? Select all that apply. - the mean of the distribution of sample means is equivalent to the population mean for larger sample sizes - the distribution of sample means is an approximately normal distribution for large sample sizes - the median of the distribution of sample means is equal to 1 - the standard deviation of the distribution of sample means equals σ / √ n - the distribution of samples means is a uniform distribution which of the following are some of the major critiques of hypothesis testing? select all that apply. - the significance level is an arbitrary value - there is publishing bias towards results that are not statistically significant - p values do not reflect the size of an effect - hypothesis testing is unaffected by sample size - the results of hypothesis testing are frequently misinterpreted and misunderstood Researchers are studying the relationship between room lighting and exam scores. After they conduct their experiment, they find a statistically significant result that provides evidence that a dimly lit room results in decreased performance on an exam. However, in reality, there is no relationship between room lighting and exam performance. What type of error was made by the researchers? Explain your answer. Type 1 error - "false positive" in this case the researchers rejected the null hypothesis when the null hypothesis was true true or false: a normal distribution follows a symmetrical, bimodal curve FALSE! a normal distribution follows a bell-shaped, symmetrical, UNIMODAL curve known as "normal curve". true or false: a normal distribution can be defined by its median FALSE! but the mean is located in the middle of the curve and the median/mean/mode are typically all equal. true or false: a standard score is an exact value that is observed false! a RAW SCORE is an exact value that is observed The ____________ shows the proportion of values that fall to the left of a given z-score in a normal distribution. - normal curve - t-table - standard normal table - normal frequency table Use the z-score to the nearest tenth the _______________ hypothesis of a z-test states that there is no difference between a given sample mean (x̄) and a population mean (μ). Whereas, the ____________ hypothesis states there is a difference. null; alternative what does a z-score of 0 indicate? the raw value corresponds to the mean in my words: the raw value is the same number as the population mean (no deviation?) what Cohen’s d value typically corresponds to a small effect size? 0.2 medium = 0.5 large = 0.8 in a normal distribution, approximately what percentage of values lie within 3 standard deviations of the mean? 99.7% - the 68/95/99.7 rule says that 68% of the values fall within 1 standard deviation of the mean, about 95% of the values fall within 2 standard deviations, and 99.7% of the values fall within 3 standard deviations. Jenna takes her male golden retriever to the vet for a yearly check-up. The dog's weight is 73 lbs. The average weight for a male golden retriever is 71 lbs with a standard deviation of 1.7 lbs. Calculate the z-score for this value. z = X - μ / σ z score = sample value - population mean / standard deviation = (73 - 71) / 1.7 = + 1.18 The customer service department of a technology company is rolling out a new protocol for handling customer issues by allowing a few select representatives to try it out first. The average customer satisfaction rating under the old protocol was 3.9 out of 5 (σ = 0.78) and the average customer satisfaction rating under the new protocol is 4.1 out of 5. Calculate the effect size. d = x̄ - μ / σ = (4.1 - 3.9) / 0.78 = 0.26, small effect? what calculation do researchers use to determine the absolute size of a treatment's effect effect size (Cohens d) Sabrina scored an 83 on both her economics and history exams. Her z-score for economics was +1.27 and her z-score for history was -1.02. On which exam did she perform better? Explain your answer. She performed better on the economics exam. The positive z-score indicates a score above the mean, whereas the negative z-score indicates a score below the mean. which type of hypothesis test is also known as a directional test? - one-tailed - two-tailed - predictive - three tailed For PE class, Jasmine's z-score for her time to run a mile was +0.7. The class average was 9.1 min with a standard deviation of 1.1 min. How much time did it actually take Jasmine to run the mile? X = x̄ + z (s) value = sample mean + z-score (standard deviation) X = 9.1 + (0.7)(1.1) = 9.87 minutes statistical power involves determining the probability of making what type of error? TYPE II ERROR (failing to reject a false null hypothesis). Statistical power (β) is effect, sample, size, significance level, number of tails of the test, and type of hypothesis used. which of the following tends to decrease statistical power? select all that apply. -large sample sizes -two tailed tests -high significance levels -one tailed tests -low α levels Note: Large effect sizes, large sample sizes, one-tailed tests, and higher significance levels tend to INCREASE statistical power. for a two tailed z-test with α = 0.05, which of the following is the correct critical value? +/- 1.96 +/- 3.30 +/- 2.575 +/- 1.86 Note: need to memorize these values, they aren't given on exam a ________ is a range of values that is likely to contain the true population mean. CONFIDENCE INTERVAL - they are centered around the mean and include a margin of error. A lab technician is testing the reliability of one of the balances in the lab. He weighs the same weigh boat with 2g of sugar 9 times and obtains a mean measurement of 2.02 g (s = 0.04). Using a critical value of 1.96, construct a 95% confidence interval for these data. confidence interval range = (x̄ - (z x s/√n) to (x̄ + (z x s/√n) (2.02 - 1.96 x 0.04 / √9) to (2.02 + 1.96 x 0.04 / √9) = 1.994 to 2.046 1.994 g < μ < 2.046 g a group of 10 local ice cream shops believes that they sell more ice cream than the average ice cream shop because they have a better-quality product than most shops. The mean number of customers they serve per day is 512. The national mean for ice cream shops is 500 with a standard deviation of 24. Perform all 4 steps of a z-test to test this claim (α = 0.05) Step 1) H0: μ = 500 Ha: μ > 500 Step 2) critical value = + 1.65 Step 3) z = (X - μ / σ) / √n = (512 - 500 / 24) / √10 = 1.58 Step 4) fail to reject null hypothesis. The z-score does not fall into the critical region. There is not enough evidence to support that there is likely a difference in the number of customers that frequent the group of ice cream shops. a teacher administers a nutrition exam to her 45 students. the mean score on the exam was 75 (μ = 75, σ = 8). The teacher wants to determine the proportion of students who scored above 80 on the exam. Using the z-table, determine this proportion. z = X - μ / σ = 80 - 75 / 8 = 0.63 the proportion of scores that lie below z = 0.63 (to the left) is 0.73565. This means that 73% of the values lie below the z-score of 0.63. To find the scores above the z-score, 0.73565 needs to be subtracted from 1. Therefore, 1 - 0.73565 = 0.26435 or 26.435% true or false: the numerator of the t-score formula contains the estimated standard error false! the DENOMINATOR contains the estimated standard error the t-distribution for a __________ sample size more closely resembles the z-distribution. - small - large - broad - general As the sample size increases, the standard deviation decreases. Therefore, with a larger sample, the t-distribution more closely resembles z-distribution. if a researcher states in the alternative hypothesis that μ is less than the claimed value, what type of test would be performed? - two tailed - left tailed - right tailed - three tailed true or false: the pooled variance is used in a one sample t test false! it is used in an INDEPENDENT samples t test true or false: when the population standard deviation is known, researchers use the t-score. false! researchers use the t-score when the population standard deviation is UNKNOWN what is the number of values that may vary freely until all remaining values are determined? - critical values - statistical power - effect size - degrees of freedom
[Show More]