Free preview

Pick a semester and a book — read its full first chapter free, no account needed.

Semester 1 Semester 2 Semester 4 Semester 8

← Back to Semester 8

Biostatistics & Research Methodology

Chapter 1: Pharmaceutical Biostatistics: Foundations and Applications

By Tapan Kumar Mahato and Satish Agrawal

Abstract

Statistics is used almost every day by everyone like calculating average which comes under measures of central tendency i.e. to get a central value. Likewise median and mode. Previous data is very helpful in future prediction.

Statistics plays a vital role here. Distribution or scattering of data is studied under measures of dispersion by means of range and standard deviation. Correlation means to find the extent of relationship between two or more variables.

This chapter explains the above topics using examples with calculations in a easy and understandable way. Pharmaceutical examples provide a better way to understand how biostatistics is beneficial in research projects. In the end a short introduction to biostatical software is given which helps students/researchers to compute fast and accurately which is not feasible by hands using pen and paper.

Keywords: Biostatistics, central tendency, mean, median, mode, dispersion, correlation. 1.1.1 Statistics: Statistics is the branch of mathematics that pertains to the study of collecting, analyzing, interpreting, presenting and organizing data in a specified mode. Data are

collected by performing surveys and experiments and by analysing the collected data by statistics to draw a conclusion from the sample data that is collected using surveys or experiments. Multiple sectors such as Medical, Pharmaceutical science, Engineering science, Psychology, Sociology, Geology and so on use statistics to function and for research. [1] Characteristics of Statistics: Some important aspects of Statistics are i. Statistics are articulated in numbers. ii.

It is gathering of facts iii. In structured arrangement, the data are collected iv. It should be analogous to each other v.

Data are gathered for a purpose which is already systematically designed. Why Statistics is important? It helps in gathering information about the relevant numerical data, reduces long complicated datasets to short graphical form, tabular form and in pictorial representation which become easy to understand, provides a well explanation which provides a clear comprehension, helps in designing the efficient and suitable planning of the statistical exploration, helps to understand the fluctuation trend through the measurable findings. [2] Statistics: In routine life [3] We cannot imagine the presence of life without statistics.

Using statistics one can sort out the problems of daily life. Statistics shapes our life and help us to understand our present, future and past. By evaluating the previously collected data using statistics, we can predict and make conclusions about the future, what will/can happen.

In our life or daily life or around us wide range of utilizations of statistics are there. Some examples are given in table 1. Table 1: Some examples of uses of statistics in Life/Daily Life S.No.

Use Explanation 1 Census Census data provides information about the population growth, religions growth, sex ratio, education level and many more. Domestic animals census is done to know the distinct traits of a specific animal like cows, dogs, buffalo, sheep, etc. 2 Sampling Blood test of any person need sample of blood for diagnosis of disease. Strength of chalk can be investigated by taking only one or two chalk from the lot. 3 Predictions Doctors use statistics to be knowledgeable about the future outlook of a certain disease e.g. control of flu outbreaks every winter by use of statistics.

Quality testing Inside the store, we try out an assortment of sweets/fruits to select the best quality option. 5 Inferences and causality Researchers collect the data from various hospitals and caregivers to determine the cause of a disease and its prevalence. 6 Business Data analysis assists a businessman in organizing production based on customer demand and satisfaction, assessing product quality, and making informed decisions about business location, marketing strategies, finance, and resources. 7 Economics The correlation between supply and demand, trade balances, inflation levels, and income per capita is determined through statistical analysis. 8 Mathematics Statistics and mathematics are mutually linked. Typically determining mean is a common thing in daily life which is a statistical tool. 9 Banking Bank examine past data and forecast the number of savings/current accounts likely to be opened next year. 10 Engineering statistics Engineering data analysis is associated with manufacturing, such as component sizes, tolerance ranges, types of materials used, and process control measures. 11 Meteorology Weather scientists predict meteorological conditions which aids in managing disasters such as floods and cyclones. This prediction is crucial for developing effective strategies to protect human and animal lives.

Agriculture Analytical methods and statistical tools are employed to solve wide range of problems linked to crop production, vegetable cultivation etc. 1.1.2 Biostatistics Biostatistics is the branch of biological science which is related to the exploration of collection, representation, evaluation, and interpretation if data of biological research. Biostatistics is also called biometrics because it requires a range of measurements and statistical calculations. In biostatistics, the statistical methods are utilized to solve biological issues.

Francos Galton is called as the father of biostatistics. For easy understanding, we can say biostatistics is combination of biology and mathematics. [4] Biostatistics’ significance in pharmaceutical research a. Designing clinical trials, which involves ensuring that studies are sufficiently robust and controlled. b.

The procedure for evaluation if a treatment is more efficient than a placebo is termed as drug efficacy analysis. c. The tracking of drug safety through the examination of undesirable outcome reports is referred to as pharmacovigilance. d. Personalised medicine describes the practice of utilizing statistical models to outlook the manner in which selected patients will have reactions to specific treatments.

Applications of biostatistics in Pharmacy a. Drug discovery: Evaluating candidate pharmaceuticals through the examination of high-throughput screening results. b. Clinical Trials: Evaluating sample sizes, assessing the results of therapy, and evaluating the safety and effectiveness of fresh medicinal compounds. c.

Epidemiology: Examining the incidence and determinants of health issues in communities. Researchers and healthcare practitioners can utilize evidence-based choices with the support of biostatistics, which confirms the correctness and stability of conclusions derived from Pharmaceutical research. [6] Application of biostatistical analysis in handling health-related uncertainties At the individual patient level, there can be possible error in judgement regarding diagnosis, treatment and health condition predictions. At the group or community level, medical uncertainty comprises a doubtfulness regarding the role of primary and short duration risk factors of many conditions of unwellness and regarding the exact effect of various health promotion, prevention and therapeutic measures.

In both these setups, the shared element is the unpredictability concerning the present state and future progression, whether with or without intervention. Example: Convalescent plasma therapy for COVID-19 patients. Here biostatistics is crucial in addressing issues in above both contexts. [7] 1.1.3 Frequency Distribution [8] In statistics, a frequency distribution is a table that illustrates the occurrence rate of different results in a sample.

Each entry in the table includes the frequency or quantifies the appearances of values within a particular group or interval. Before constructing a frequency table, it is beneficial to understand about the range (minimum and maximum values). The range is segmented into arbitrary intervals called “class interval.” The data can be showcased either in a graphical or tabular form that shows the count of observations within a given interval.

The intervals should be distinct and comprehensive. 1.1.4 Pharmaceutical examples Example 1: The following is the distribution of the weight of 150 students (Girls and boys both) in a Pharmacy college. Determine the lower limit of the first class interval, the class limits of the third class, the class mark for the interval 45-50 and the class size. Table 2: Frequency distribution table of weight of 150 students (Example 1) Weight No. of students 40-45 30 45-50 45 50-55 50 55-60 25 Total 150 Solution: The lower limit of the first-class interval (40-45) is 40, the class limit of the third class (50-55) are 50 (lower limit) and 55 (upper limit), the class mark is the average of the class limits of a class i.e. for 45-50, the class mark is (45 + 50)/2 = 47.5 and the class size is the difference between the lower and upper class-limits i.e. 5 (differences between 45 - 40, 50 - 45, 55 - 50, 60 - 55 are all equal to 5).

Example 2: In a hospital, the recorded ages of 48 patients visited OPD on 01 January 2025 are given below. Construct a frequency distribution table for the provided data. Table 3: Ages of 48 patients (Example 2) 5 8 2 0 5 6 3 7 2 4 3 5 2 3 1 4 2 4 4 0 5 1 4 4 5 9 6 2 1 7 2 5 4 2 3 5 1 7 2 6 5 8 6 2 4 5 2 4 3 2 4 5 4 1 4 7 3 4 6 3 2 3 3 8 1 5 2 9 3 6 5 2 1 4 1 5 5 9 5 7 2 5 2 7 1 8 1 9 1 1 1 3 2 3 6 4 Solution: Arranging the provided ages in ascending order, we get 11, 13, 14, 14, 15, 15, 17, 17, 18, 19, 20, 23, 23, 23, 24, 24, 24, 25, 26, 27, 27, 29, 32, 34, 35, 35, 36, 37, 38, 40, 41, 42, 44, 45, 45, 47, 51, 52, 56, 57, 58, 58, 59, 59, 62, 62, 63, 64.

Table 4: Frequency distribution table (Example 2) Class Interval Frequency 10-20 11 20-30 11 30-40 08 40-50 06 50-60 08 60-70 04 Total 48 Example 3: A survey was carried out in 22 medical stores (M1 to M22) to find selling of paracetamol tablets packets of 500 mg (each packet contains 10 strips of 100 tablets) in a day. The results were recorded is given in table 5 below. Prepare a Frequency Distribution Table from the given data.

Table 5: Selling of paracetamol tablets in 22 medical stores (Example 3) M 1 M 2 M 3 M 4 M 5 M 6 M 7 M 8 M 9 M 1 0 M 1 1 M 1 2 M 1 3 M 1 4 M 1 5 M 1 6 M 1 7 M 1 8 M 1 9 M 2 0 M 2 1 M 2 2 3 1 7 4 0 2 6 8 1 5 2 1 5 4 2 3 7 2 0 9 2 1 Solution: Arranging the provided data in ascending order, we get 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 2, 3, 3, 4, 4, 5, 5, 6, 7, 7, 8, 9 Table 6: Frequency Distribution Table (Example 3) Number of packets Frequency 0 2 1 4 2 5 3 2 4 2 5 2 6 1 7 2 8 1 9 1 Total 22 1.2 Measure of central tendency 1.2.1 Introduction: Central tendency is a statistical measure that identifies/determines a single value as representative of an whole dataset. It gives an accurate summary of the data by targeting on the centre value around which the data points are gathered. The three ways used to measure of central tendency are: i.

Mean: The average of the dataset. ii. Median: The middle value after arranging the data in ascending or descending order. iii. Mode: The value that appears most frequently or repeatedly in the dataset. 1.2.2 Mean The arithmetic mean or mean or average is a measure of central tendency that is determined by adding all the values in a dataset and dividing up by the number of values. [9] The mean is a basic metric in statistics that illustrates the central value of a dataset.

The mean provides a single value that summarizes the data, offering insights into the dataset's central tendency. 1.2.3 Types of Mean i. Arithmetic Mean: The most commonly used type of mean. It is calculated by dividing the sum of all observations by the number of observations. ii.

Geometric Mean: Used for datasets involving rates of change or multiplicative factors. Obtained by extracting the nth root from the overall product of observations. iii. Harmonic Mean: Appropriate for datasets where the average of rates is desired.

Calculated by dividing the count of observations by the aggregate of the reciprocals of the observations. These different types of means are utilized based on the nature of the data and the specific requirements of the analysis. [10] 1.2.4 Methods to Calculate Mean a. Direct method b.

Assumed Mean Method or Shortcut method c. Step deviation method a. Direct Method Used for smaller datasets or when raw data is provided.

Arithmetic Mean = x 1 + x 2 +…………+ x n = ∑x -------------------------------------- ------------- N N Where, x 1 to x n = Individual observations, ∑x = Sum of all observations and N = Total number of observations b. Assumed Mean Method (Shortcut method) Used for larger datasets to simplify calculations. Where, A: Assumed Mean, f i : Frequency of the i-th class, d i =x i −A: Deviation of the class midpoint (x i ) from the assumed mean. c.

Step-Deviation Method Useful for grouped data or large datasets with uniform intervals. ∑fu Arithmetic Mean = A + h x ------------ ∑f Where, x−A A = Assumed mean, u = Step deviations (u = ---------), h h = Class width (common factor) [11] 1.2.5 Pharmaceutical examples to calculate mean Example 4: The marks obtained by 10 students in a class test of pharmaceutical chemistry subject are 60, 56, 44, 20, 50, 80, 78, 64, 94, 72 out of 100. Find the arithmetic mean (AM) by direct method. Solution: Here, the marks obtained are 60, 56, 44, 20, 50, 80, 78, 64, 94, 72.

Number of students (N) = 10 Therefore, the calculation is as follows, Sum of marks (x 1 +x 2 +x 3 ….. + x n ) ∑x AM = --------------------------------------------- = -------- Number of students N 60 + 56 + 44 + 20 + 50 + 80 + 78 + 64 + 94 + 72 618 AM = --------------------------------------------------------------------- = -------------- = 61.8 10 10 Calculation of arithmetic mean for grouped data Example 5: The following table represents the distribution of marks scored by B.Pharm. first semester students in a class test of remedial biology subject: Table 7: distribution of marks of 100 students (Example 5) Marks Frequency 00-10 05 10-20 10 20-30 20 30-40 30 40-50 25 50-60 10 Calculate the Arithmetic Mean of the marks using the Assumed Mean Method. Solution: 1. Calculate Class Midpoints (x i ): 2.

Choose an Assumed Mean (A): Let’s select A = 35 (the midpoint of the interval 30–40). 3. Compute Deviation (d i = x i − A): Find the deviation of each midpoint from assumed mean. 4. Multiply Frequency by Deviation (f i ⋅ d i ). 5.

Apply the Formula: Table 8: Shows table of marks, frequency, mid points, assumed mean, deviation and product of frequency and deviation (Example 5) Marks Frequency (f i ) Mid point (x i ) Assumed mean (A) Deviation ( d i =x i −A ) f i .d i 00-10 05 05 35 -30 -150 10-20 10 15 35 -20 -200 20-30 20 25 -10 -200 30-40 30 35 00 000 40-50 25 45 10 250 50-60 10 55 20 200 Total ∑f i = 100 ∑f i d i = -100 Applying the formula (-100) AM = 35 + ------------- = 35 - 1 = 34 100 Calculation of Arithmetic Mean Using the Step-Deviation Method The step-deviation method is a streamlined approach to calculate the arithmetic mean, especially useful when dealing with large datasets or when the data values are large. This method makes calculations easy and simple by decreasing the size of the numbers involved in the dataset. Example 6: Consider the following frequency distribution of marks obtained by students: Table 9: Marks distribution of 36 students (Example 6) Marks interval Frequency 0-10 8 10-20 5 20-30 10 30-40 6 40-50 7 Solution: Steps to calculate the Arithmetic mean using the Step-Deviation method: 1.

Assumed Mean (A): Select a value within the data range as the assumed mean. 2. Class Interval (h): Determine the width of each class interval. 3. Midpoint (x): Calculate the midpoint for each class interval. 4.

Deviation (d): Compute the deviation of each midpoint from the assumed mean: d=x−A. 5. Step-Deviation (u): Divide each deviation by the class interval width: u=d/h . 6. Frequency (f): Note the frequency of each class interval. 7.

Calculate f×u: Multiply the frequency by the step-deviation for each class. 8. Summation: Sum up the frequencies (∑f) and the products (∑f×u). 9. Arithmetic Mean Formula: Apply the formula: Table 10: Shows marks, class interval, frequency, mid points, assumed mean, deviation, step deviation and product of frequency & step deviation (Example 6) Marks interval Class interval (h) Frequency (f) Mid point (x) Assumed mean (A) Deviation (d) Step deviation (u) fu 0-10 10 8 5 25 -20 -20/10 = -2 -16 10-20 10 5 15 25 -10 -10/10 = -1 -05 20-30 10 10 25 25 0 00/10- 00 00 30-40 10 6 35 25 10 10/10 = 1 06 40-50 10 7 45 25 20 20/10 = 2 14 Total ∑f = 36 ∑fu = -01 On substituting values in the formula, we get -1 = 25 + 10 x ------------------------ = 25 + 10 x (-0.028) = 25 - 0.28 = 24.72 Therefore, the arithmetic mean is 24.72. [12] 1.2.6 Geometric Mean The geometric mean is calculated by determining the n th root of the multiplication of all observations.

It is suitable for data involving ratios, percentages, or growth rates. Formula for Ungrouped Data: Geometric Mean = (Product of all observations) 1/N = P 1/N Where, x : Individual observations, N: Total number of observations and P: Product of all observations Formula for Grouped Data: Where, f: Frequency of the observation, x: Midpoint of the class interval and ∑f: Total frequency. [13] 1.2.7 Pharmaceutical examples to calculate geometric mean Example 7: The geometric mean (GM) of ten values: 2, 3, 4, 5, 6, 1, 7, 8, 9, 10 is: (2 x 3 x 4 x 5 x 6 x 1 x 7 x 8 x 9 x 10) 1/5 GM= 5 √3628800 = 20.509 Example 8: The table below shows the frequency distribution of the wages (in rupees) of workers in a drug manufacturing company named Galaxy pharmaceutical Pvt. Ltd.

Calculate the geometric mean (GM) of the wages of the workers. Table 11: distribution of eages of 33 workers (Example 8) Wages 10-20 20-30 30-40 40-50 50-60 Frequency 4 6 10 8 5 Solution Step 1: Find out the Midpoints (x) Step 2: from the obtained midpoints, calculate ln x (natural logarithm). Step 3: find out f.ln x by multiplying f by ln x Step 4: Find ∑f and total ∑f.ln x Step 5: Apply the Formula and get the geometric mean Lower Limit + Upper Limit x = ------------------------------------------ 2 Table 12: Shows wages, mid points, frequency, ln x and product of frequency & ln x (Example 8) Wages Midpoint (x) Frequency (f) ln x f.ln x 10-20 15 4 2.708 10.832 20-30 25 6 3.219 19.314 30-40 35 10 3.555 35.550 40-50 45 8 3.807 30.456 50-60 55 5 4.007 20.035 Total ∑f = 33 ∑f.ln x = 116.187 The formula to calculate geometric mean is: On substituting the values, we get GM = exp (116.187 / 33) = exp (3.521) = 33.93 1.2.8 Harmonic Mean The harmonic mean (HM) is a statistical measure used primarily when dealing with rates or ratios.

It is the reciprocal of the arithmetic mean of reciprocals of a dataset. The use of harmonic mean is most appropriate when data involves rates (e.g., speed, efficiency) and when averaging quantities like distances or times. The formula used to calculate harmonic mean is: Where, N: Total number of observations and x i : Individual observations For ungrouped data, the formula becomes: Where, f is frequency of each observation and x is the midpoint of class intervals. [14] 1.2.9 Pharmaceutical examples to calculate harmonic mean Example 9: Find out the harmonic mean of the rejected tablets in manufacturing of diclofenac tablets in six batches: 60, 4, 36, 45, 50, 75 is (here N = 6 and x 1 = 60, x 2 = 4, x 3 = 36, x 4 = 45, x 5 = 50, x 6 = 75). 6 6 HM = ---------------------------------------------- = ------------------------------------------------------- 1/60 + ¼ + 1/36 +1/45 +1/50 +1/75 0.017 + 0.25 + 0.028 + 0.022 + 0.02 + 0.013 6 HM = ------------------- = 2.85 0.35 Calculation of harmonic mean of grouped data For grouped data, the harmonic mean can be calculated using the following formula: Where, N = Total number of observations, f i = Frequency of the i th class and x i = Midpoint of the i th class.

Example 10: Consider the following frequency distribution representing the speeds (km/h) of ambulances carrying patients in emergency cases observed on a highway during peak hours of traffic. Table 13: Shows speeds of 25 ambulances (Example 10) Speed (km/h) Frequency (f i ) 20-30 5 30-40 8 40-50 12 50-60 7 60-70 3 Solution: 1. Calculate the midpoints (x i ) of each class interval. 2.

Compute f i /x i for each class. 3. Find out ∑fi/xi 4. Total frequency (N). 5.

Calculate the Harmonic Mean (H.M.) Table 14: Shows speed, frequency, mid points and division of frequency by mid points (Example 10) Speed (km/h) Frequency (f i ) x i f i /x i 20-30 5 25 0.20 30-40 8 35 0.23 40-50 12 45 0.27 50-60 7 55 0.13 60-70 3 65 0.05 Total ∑f i = N = 35 ∑f i /x i = 0.88 35 HM = ------------------------- = 39.77 0.88 Answer: The harmonic mean of the ambulance speeds is approximately 39.77 km/h. [15] 1.2.10 Median The median is a measure of central tendency that tells about the middle value of a dataset when the values are organized in increasing or decreasing sequence. The median divides a dataset into two equal segments. Calculation of median for grouped or ungrouped data: 1.

For Ungrouped Data i. If N (the total number of observations) is odd then Median = Middle Value ii. If N is even then 2.

For Grouped Data The median is determined by applying the formula: Where, L: Lower boundary of the median class, N: Total frequency, CF: Cumulative frequency before the median class, f: Frequency of the median class and h: Width of the median class Steps to Calculate Median for Grouped Data 1. Prepare a cumulative frequency table. 2. Find N/2 where N is the total frequency. 3.

Identify the median class (the class containing N/2. 4. Apply the formula to calculate the median. [16] 1.2.11 Pharmaceutical examples to calculate median Example 11: The individual weight (in mg) of 15 sample tablets are given below. Find the median weight of the tablets. 10, 27, 11, 23, 37, 14, 21, 21, 11, 22, 26, 17, 13, 15, 20 Solution: Let us arrange the obtained weight of tablets in ascending order: 10, 11, 11, 13, 14, 15, 17, 18, 21, 21, 22, 23, 26, 27, 37 Here, number of observations = 15 (odd) 10, 11, 11, 13, 14, 15, 17, median value, 21, 21, 22, 23, 26, 27, 37 Answer : The median will be the mid value i.e. the 8th observation which is 18 mg.

Example 12: Below is a frequency distribution of marks obtained by M.Pharm. first year students in a class test of quality assurance and quality control subject. Calculate the median of the marks obtained by students. Table 15: Marks distribution of 35 students (Example 12) Marks range 0-10 10-20 20-30 30-40 40-50 Frequency (f) 5 8 10 7 5 Solution Table 16: Shows marks range, frequency, cumulative frequency and median class (Example 12) Marks range Frequency (f) Cumulative frequency (CF) N/2 Median class 0-10 5 5 35/2 = 17.5 10-20 8 13 20-30 10 23 20-30 30-40 7 30 40-50 5 35 Total ∑f = N = 35 The median class is the class where the cumulative frequency first exceeds 17.5.

From the table, this is 20 - 30. By inserting the values into the formula, we obtain N/2 – 13 17.5 – 13 4.5 Median = 20 + ---------------------------- x 10 = 20 + ---------------- x 10 = 20 + -------- x 10 10 10 10 = 20 + 4.5 = 24.5 Answer: The median of marks of students obtained in class test is 24.5. 1.2.12 Mode The mode is a measure of central tendency. The most frequently appearing value in a dataset is mode.

It is especially advantageous in datasets where the most repeating item, event, or value. Calculation of mode for ungrouped data The mode is the value that appears most repeatedly. If all values appear with the identical frequency, the dataset has no mode.

A dataset may additionally possess more than one mode (bimodal or multimodal). Calculation of mode for grouped data: The mode is determined using the following formula: Where, L: Lower boundary of the modal class, f 1 : Frequency of the modal class, f m : Frequency of the class preceding the modal class, f 2 : Frequency of the class succeeding the modal class and h: Class width Steps to Calculate Mode for Grouped Data 1. Find out the modal class i.e. class interval with the highest frequency. 2.

Find out the frequencies of the modal class (f 1 ), the class before it (f m ), and the class after it (f 2 ). 3. Substitute the values in the formula to calculate the mode. [17] 1.2.13 Pharmaceutical examples to calculate mode Example 13: In the data set of selling of cough syrup bottles of 100 ml in 10 medical stores on 30 January 2025 are 3, 5, 1, 2, 4, 9, 4, 5, 7, 5, the mode is 5 because it appears more often than any other number. Example 14: For categorical data, in a pen set of 10.

The colours of the pens are red, blue, blue, green, white, black, yellow, blue, red, pink, the mode is "blue" because it occurs maximum times. Example 15: The following table shows the distribution of heights (in cm) of 60 students in a B.Pharm. first year class in a pharmacy college: Table 17: Shows heights of 60 students (Example 15) Height (cm) 140-150 150-160 160-170 170-180 180-190 Frequency (f) 7 10 22 12 9 Find the mode of the distribution. Solution: Step 1: The modal class is the class with the highest frequency.

Here, the highest frequency is f m = 20, corresponding to the class 160–170. So, the modal class is 160–170. Step 2: l = 160, f m = 22, f 1 = 10 (Frequency of the class previous to the modal class =150–160), f 2 = 12 (Frequency of the class next to the modal class = 170–180) and h = 10 (Class width (difference between upper and lower borders of any class).

Step 3: 22 – 10 12 Mode = 160 + ------------------------------- x 10 = 160 + ---------------- x 10 (2 x 22) – 10 - 12 44 -10-12 12 = 160 + ---------- x 10 = 160 + (120 / 22) = 160 + 5.45 = 165.45 22 Answer: the mode of heights of 60 students is 165.45 cm. [18] 1.2.14 Pharmaceutical examples to calculate mean median and mode Example 16: Find mean, median and mode of tablet hardness data obtained for a tablet production batch. The obtained data is 3, 6, 4, 9, 13, 10 and 6. Solution: Sum of all data 3 + 6 + 4 + 9 + 13 + 10 + 6 51 Mean =------------------------------------------ = ------------------------------------- = --------- = 7.29 Number of data 7 7 Median - For median, arrange the data first 3, 6, 4, 9, 13, 10, 6 in ascending order, we get 3, 4, 6, 6, 9, 10, 13.

The middle term will be the median i.e. 6. Mode – In the data set, 3, 6, 4, 9, 13, 10, 6 the maximum appearing term is 6 is the mode because this term is appearing two times, others only one time. Example 17: The relative humidity of six days (Monday to Saturday) in the liquid section of a pharmaceutical company is given in table below.

Calculate the mean percent relative humidity. Table 18: Shows percent relative humidity of six days in a week from Monday to Saturday (Example 17) Day Percent relative humidity Monday 52 Tuesday 54 Wednesday 57 Thursday 61 Friday 67 Saturday 57 Solution: - ∑ x 52 + 54 + 57 + 61 + 67 + 57 348 x = ---------- = ------------------------------------------ = ------------ = 58 N 6 6 Example 18: Seven people came to hospital in first hour on Monday for COVID 19 Vaccination in the month of september 2020. The ages (in years) of those people are: 56, 67, 54, 34, 78, 43, 23.

What will be the median? Solution: On arranging the data in ascending order, we get: 24, 35, 44, 55, 57, 68, 79. Here, the number of observations = 7 (odd number) 24, 35, 44, Median value, 57, 68, 79 Median value = 4th observation Median = 55 Answer: The median of ages of the people is 55 years.

Example 19: Six people came to hospital in second hour on Monday for COVID 19 Vaccination in the month of september 2020. The ages (in years) of those people are: 51, 68, 25, 35, 79, 44. What will be the median?

Solution: On arranging the data in sequence of smaller to larger then we get: 25, 35, 44, 51, 68, 79. Here, the number of observations = 6 (even number) 6/2 = 3, now using the formula, Median = (3rd obs. + 4th obs.) / 2 = (44 + 51) / 2 Median = 47.5 Answer: The median of the ages of the people is 47.5 years. Example 20: Find the mode of the given data of admitted patients in orthopaedic IPD in a hospital: Table 19: Distribution of ages of 36 patients (Example 20) Age (years) Number of patients 00-20 03 20-40 06 40-60 12 60-80 10 80-100 05 Solution: The highest frequency = 12, so the modal class is 40-60. l = 40 (Lower limit of modal class) f m = 12 (Frequency of modal class) f 1 = 6 (Frequency of class preceding modal class) f 2 = 10 (Frequency of class succeeding modal class) h = 20 (Class width) on inserting the values in formula, = 40 + [12 – 6 / 2 x 12 – 6 – 10] x 20 = 40 + (6/8) x 20 Mode = 40 + 0.75 x 20 = 40 + 15 = 55 Answer: the mode of admitted patients is 55 years.

Example 21: Calculate the mean for a dataset if mode and median are given i.e. 50 and 30 respectively. Solution: If mode and median are given then we can find the mean using the formula 2 Mean + Mode = 3 Median or 2 Mean = 3 Median - Mode Then, 2 Mean = 3 x 30 – 50 2 Mean = 90 – 50 = 40 Mean = 40/2 = 20 Answer: The mean is 20 (when mode and median are given) 1.3 Measures of dispersion 1.3.1 Introduction: Measures of dispersion, also known as measures of variability or scatter, determines the extent to which data points in a dataset differ from the central value in numbers, providing understanding into the distribution of the data. Common measures include: i.

Range : in the dataset, the difference between the maximum and minimum values. ii. Variance : The average of the squared differences between each data point and the mean, indicating how data points scatter from the mean. iii. Standard Deviation : The square root of the variance, representing the average distance of each data point from the mean. [19] 1.3.2 Dispersion Dispersion refers to the extent to which data points in a dataset are dispersed around a central value, such as the mean or median.

It provides the degree of variability or diversity within a dataset as a numerical value and helps in assessing the dependability and stability of the data. [20] 1.3.3 Range Example 22 : Let’s take the dataset: 6, 8, 9, 10, 12 Calculation of Range : Here, 12 is the maximum value and 6 is the minimum value. Subtracting both values gives range. Range = Maximum value - Minimum value = 12 - 6 = 6 Answer: The range is 6 for the above data set.

To calculate standard deviation arithmetic mean and variance are also required along with range. Calculation of Mean (µ) : µ = Sum of observations / Number of observations = (6 + 8 + 9 + 10 + 12) / 5 = 45 / 5 = 9 Answer: The arithmetic mean is 9 for the above data set. Variance (σ²) : σ² = [Summation of (Individual observation – Mean) 2 ] / Number of observations σ² = [(x 1 -µ) 2 + (x 2 -µ) 2 + (x 3 -µ) 2 + (x 4 -µ) 2 + (x 5 -µ) 2 ] / Number of observations σ² = [(6 - 9)² + (8 - 9)² + (9 - 9)² + (10 - 9)² + (12 - 9)²] / 5 σ²= [9 + 1 + 0 + 1 + 9] / 5 σ² = 20 / 5 or σ² = 4 Answer: The variance of the given data set is 4. 1.3.4 Standard Deviation (σ) : σ = square root of variance = √σ² σ = √4 ≈ 2 Answer: The standard deviation of the given data set is 2. [21] 1.3.5 Pharmaceutical examples for calculation of range, variance and standard deviation of grouped data Example 23: In a week from 01/01/2025 to 07/01/2025, the sale of Ranitidine tablets 150 mg in seven medical stores of Raipur city is given below.

Calculate the mean, variance, and standard deviation for the following distribution: Table 21: Shows class interval, frequency, mid points, product of frequency & mid points (Example 23) Class interval Frequency 30-40 3 40-50 7 50-60 12 60-70 15 70-80 8 80-90 3 90-100 2 Solution: Calculate the midpoints (x i ) of each class interval then compute f i for each class then construct the following table to find out other quantities then calculate the mean (xˉ) by formula And then calculate the variance (σ 2 ) and standard deviation (σ) Table 22: Calculation of x i , x i -Mean, x i -Mean) 2 and f i .(x i -Mean) 2 Class interval Frequency (f i ) Mid point (x i ) f i x i x i -Mean (x i -Mean ) 2 f i .(x i -M ean) 2 30-40 3 35 105 -27 729 2187 40-50 7 45 315 -17 289 2023 50-60 12 55 660 -7 49 588 60-70 15 65 975 3 9 135 70-80 8 75 600 13 169 1352 80-90 3 85 255 23 529 1587 90-100 2 95 190 33 1089 2178 Total 50 3100 10050 Step 1: Calculation of (xˉ) Mean xˉ = 3100 / 50 = 62 Answer: Arithmetic mean of the given data is 62. Step 2: Calculation of Variance (σ 2 ) 10050 = -------------------------------- = 201 50 Answer: the variance of given data is 201. Step 3: Calculation of Standard Deviation (σ) = √201 = 14.177 Answer: The standard deviation of the given data is 14.177 Step 4: Calculation of Range Range = Maximum value − Minimum value = 100 – 30 = 70 Answer: The range of the given data is 70. [22] 1.4 Correlation 1.4.1 Definition and concept of correlation Correlation defines the magnitude and tendency of the relationship between two variables.

It tells in numerical value the change in one variable associated with change in another. The correlation coefficient is denoted as r whose value ranges from -1 to +1. Let us discuss the meaning of range from -1 to +1.

  • +1 indicates a perfect positive linear relationship i.e. one variable increases, the other also increases proportionally.
  • 0 (zero) indicates that there is no linear relationship between the variables.
  • -1 indicates that a perfect negative linear relationship is there i.e. one variable increases, the other variable decreases proportionally. [23] 1.4.2 Karl pearson’s coefficient of correlation Karl Pearson's coefficient of correlation, denoted by r is a measure of the linear relationship between two variables. It is used to quantify the strength and direction of the association between two variables. The value of r ranges from −1 to +1, which means o If r = +1, then it indicates that a perfect positive correlation is there. o If r = −1, then it indicates that a perfect negative correlation is there. o If r = 0, then it indicates that no linear correlation is there. It is non-dimensional i.e. it is

not affected by the units of the variables. The Karl Pearson's coefficient of correlation is applicable to fields like biology, finance, science and social sciences to determine if one variable can predict another in regression analysis. Formula to calculate the Karl Pearson's coefficient of correlation (r) Where, x i and y i are the data points for the two variables X and Y and xˉ and yˉ are the means of X and Y. 1.4.3 Numerical problem to calculate the Karl pearson’s coefficient of correlation Example 24: Calculate karl pearson’s coefficient of correlation (r) for the following dataset: Table 23: Dataset for calculation of Karl Pearson’s coefficient of correlation (Example 24) X 1 2 3 4 Y 2 3 6 8 Solution: Steps to follow to solve the problem are first, summarize the data i.e. computation of ∑X, ∑Y, ∑X 2 , ∑Y 2 and ∑(XY).

Second, use the formula and third, insert the values in the formula and solve. Table 24: Shows values of X, Y, X 2 , Y 2 and XY (Example 24) X Y X 2 Y 2 XY 1 2 1 4 2 2 3 4 9 6 3 6 9 36 18 4 8 16 64 32 ∑X = 10 ∑Y = 19 ∑X 2 = 30 ∑Y 2 = 113 ∑(XY) = 58 Using the formula: On substituting the values in the formula, we get: – 190 42 r = ----------------------------------------- = ------------------------------------ √ (120 - 100) (452 - 361) √ 20 x 91 42 42 r = --------------------------------------- = -------------------------------- = 0.984 √ 1820 42.66 Answer: The coefficient of correlation is r ≈ 0.984 which indicates a very strong positive linear relationship between X and Y. [24] 1.4.4 Multiple correlation Multiple correlation tells the relationship between one dependent variable and two or more independent variables. It determines how well the dependent variable is predicted by a combination of independent variables.

The multiple correlation is used to study complicated systems incorporating several variables and also applied in predictive modelling like assessing drug effectiveness based on multiple factors (dosage, patient age, weight and genetic profile). The multiple correlation coefficient (R) is a statistical measure that represents the degree of relationship between a dependent variable (Y) and two or more independent variables (X 1 , X 2 , X k ). The value of multiple correlation ranges from 0 to 1.

  • If R = 0, it indicates that there is no relationship between the dependent variable and the set of independent variables.
  • If R=1, it indicates that there is perfect prediction of the dependent variable by the independent variables. Formula to calculate multiple correlation: For two independent variables (X 1 and X 2 ) predicting Y, the multiple correlation coefficient R can be calculated as: Alternatively, R 2 , the coefficient of determination, shows the proportion of variance in Y explained by the independent variables. The key concepts to keep in mind are: a. Multiple Regression: Used to compute the regression equation: Y=a + b 1 X 1 + b 2 X 2 + ⋯ + b k X k Where, a: Intercept, b 1 , b 2 , … , b k are regression coefficients for X 1 , X 2

, … , X k b. Interpreting R: o R quantifies how well the independent variables collectively predict Y. o Higher R values indicate stronger relationships. c. Adjusted R 2 : Adjusted R 2 justifies the number of predictors in the model, providing a more accurate measure of suitability. 1.4.5 Numerical problem showing calculation of multiple regression Example 25: Suppose three quantities are there and we want set up a relation among these three e.g. to predict a dependent variable Y (blood pressure) depended on two independent variables X 1 (age) and X 2 (Weight).

Calculate the multiple correlation coefficient (R) and derive the regression equation: Y = a + b 1 X 1 + b 2 X 2. Table 25: Dataset of blood pressure, age and weight (Example 25) Y (Blood pressure) X 1 (Age) X 2 (Weight) 120 25 65 130 35 70 140 45 75 150 55 80 Solution: Steps to solve the problem are first, summarize the data. Second, compute the following for each variable i.e. ∑Y, ∑X 1 , ∑X 2 , ∑Y 2 , ∑X 1 2 , ∑X 2 2 , ∑YX 1 , ∑YX 2 , ∑X 1 X 2 and third, compute the regression coefficients using the formulas: 1.

Calculation of R 2 (Coefficient of Determination) 2. Calculation of R: Table 26: Shows values of Y, X 1 , X 2 , Y 2 , X 1 2 , X 2 2 , YX 1 , YX 2 and X 1 X 2 (Example 25) Y X 1 (Age) X 2 (Weight) Y 2 X 1 2 X 2 2 YX 1 YX 2 X 1 X 2 120 25 65 14400 625 4225 3000 7800 1625 130 35 70 16900 1225 4900 4550 9100 2450 140 45 75 19600 2025 5625 6300 10500 3375 150 55 80 22500 3025 6400 8250 12000 4400 ∑Y = 540 ∑X 1 = 160 ∑X 2 = 290 ∑Y 2 = 73400 ∑X 1 2 = 6900 ∑X 2 2 = 21150 ∑YX 1 = 22100 ∑YX 2 = 39400 ∑X 1 X 2 = 11850 Substituting the values in the formulas, we get [22100 – (160x 540)/4] [22100 – (86400/4)] b 1 = -------------------------------------------- = ……………………………… [6900 – (160) 2 /4] [6900 – 25600/4] 22100 – 21600 500 b 1 = -------------------------------- = -------------------- = 1 6900 - 6400 500 [39400 – (290x 540)/4] [39400 – (156600/4)] b 2 = -------------------------------------------- = ……………………………… [21150 – (290) 2 /4] [21150 – 84100/4] 39400 – 39150 250 b 2 = -------------------------------- = -------------------- = 2 21150 - 21025 125 1. Formation of regression equation: Y = a + b 1 X 1 + b 2 X 2 (Substituting values in this equation) Y = −50 + 1 ⋅ X 1 + 2 ⋅ X 2 Where, Intercept (a) = −50, b 1 (Coefficient for X 1 ) = 1 and b 2 (Coefficient for X 2 ) = 2. 2.

Calculation of coefficient of determination (R 2 ): R 2 = 2 This result suggests an overestimated explained variance due to data conditions. 3. Calculation of multiple correlation coefficient (R): R = √R 2 =1.414 Interpretation:

  • The regression equation provides a model to predict Y based on X 1 (Age) and X 2 (Weight).
  • The computed R 2 and R values suggest possible issues (e.g., multicollinearity) or an unrealistic dataset. Proper adjustments or checks on data quality are necessary for practical cases. [25] 1.4.6 STATISTICAL SOFTWARE Calculations using pen and paper needs labour and time to complete. To overcome this problem, several statistical programmes are available e.g. SPSS, R Online and Minitab, Microsoft Excel. Out of these Excel is used mostly because it is free and simple to use. If a person how to handle the software, he/she can complete his/her work in a very less time with accuracy. Microsoft excel: Excel is a statistical software used widely for fundamental data handling activities like tabulating data, making pivot tables, and making graphics. Excel is

part of the Microsoft Office suite. It is widely used to calculate central tendency i.e. mean, median and mode, dispersion i.e. standard deviation and correlation because it is free and easy to use that saves time, labour and money. SPSS (Statistical Package for Social Sciences): In a very less time, this software can change and analyse large amount of data in such a way that would be very difficult to execute by hands using pen and paper.

It was made by IBM. SPSS take data and convert them to tabulated reports, charts and plots of distributions and can perform many numerous statistical analysis. Minitab: Gives analysis results more precise.

It aids in exploratory data analysis by using sophisticated graphs, charts, and other tools. R: R is a component of open-source software and a programming language. R is a programming language mostly used for enhancement of statistical software along with data analysis. [26, 27] Conclusion: This chapter introduces Statistics and Biostatistics, which are essential tools for research and projects.

Understanding these concepts helps in many ways, such as predicting success rates, performing quick and accurate calculations, and simplifying large datasets into smaller, more manageable forms. For example, a frequency distribution table can summarize large amounts of data into a concise and easy-to-understand format. The chapter also explains central tendency, which includes measures like mean, median, and mode.

These are explained with simple examples from the pharmaceutical field to make them easier to grasp. Additionally, measures of dispersion, such as range and standard deviation, are discussed with numerical problems related to pharmacy. The concept of correlation is also covered, which shows how two or more sets of data are related.

Methods like Karl Pearson’s coefficient of correlation and multiple correlation are explained using pharmaceutical examples. Finally, the chapter introduces statistical software like MS-Excel, SPSS, Minitab ® , and R - Online Statistical Software. These tools are incredibly useful because they save time and provide highly accurate results compared to manual calculations.

They are particularly valuable in industrial and clinical trials, where large datasets and complex analyses are common. DISCLAMER: Assistance from ChatGPT AI, Deepseek AI, Quillbot Plagiarism checker and Galaxy synonym generator were utilized in drafting this chapter. REFERENCES 1.

Cuemath. Statistics: Definition, examples, mathematical statistics. Cuemath.

Retrieved January 29, 2025, from https://www.cuemath.com 2. BYJU'S. Statistics: Definition, types, importance, and examples.

BYJU'S. Retrieved January 29, 2025, from https://www.byjus.com 3. Singh, A.

(2019). Statistics: In routine life. International Journal of Educational Science and Research, 9 (4), 7–14. 4.

EasyBiologyClass. Introduction to biostatistics: Applications of biostatistics. EasyBiologyClass.

Retrieved January 29, 2025, from https://www.easybiologyclass.com 5. Rosner, B. (2015).

Fundamentals of biostatistics (8th ed.). Brooks/Cole, Cengage Learning. 6. Mahato, T.

K., Patel, A. P., Sharma, S., & Kushwaha, R. S.

(2024). Biostatistical methodologies in epidemiology: Importance, distinction, and computation of relative risk and attributable risk. International Journal of Pharmaceutical Science and Research, 9 (4), 9–14. 7.

Indrayan, A. (2021). Medical biostatistics as a science of managing medical uncertainties.

Indian Journal of Community Medicine, 46 , 182–185. 8. Singh, G., & Mahato, T. K.

(2022). Biostatistics & research methodology (1st ed.). Technical Publications. 9.

Newbold, P., Carlson, W. L., & Thorne, B. (2013).

Statistics for business and economics (8th ed.). Pearson Education. 10. Testbook.com.

Mean in math: Meaning, statistics along with types & formulas. Testbook. Retrieved January 29, 2025, from https://www.testbook.com 11.

Gupta, S. C. (2018).

Fundamentals of statistics (7th ed.). Himalaya Publishing House. 12. Gupta, S.

P. (2018). Measures of central tendency.

In Statistical methods (44th ed., Chapter). Sultan Chand & Sons. 13. Levin, R.

I., & Rubin, D. S. (2013).

Statistics for management (7th ed.). Pearson Education. 14. Gupta, S.

C., & Kapoor, V. K. (2020).

Mathematical statistics (8th ed.). Sultan Chand & Sons. 15. National Council of Educational Research and Training (NCERT).

(2005). Statistics for economics (Class XI). NCERT. 16.

Mood, A. M., Graybill, F. A., & Boes, D.

C. (1973). Introduction to the theory of statistics (3rd ed.).

McGraw Hill Education. 17. Gupta, S. C., & Kapoor, V.

K. (2022). Mathematical statistics (12th ed.).

Sultan Chand & Sons. 18. National Council of Educational Research and Training (NCERT). (2024).

Statistics. In Mathematics textbook for class 13 (Chapter). NCERT. 19.

Wikipedia contributors. Statistical dispersion . Wikipedia, The Free Encyclopedia.

Retrieved January 29, 2025, from https://en.wikipedia.org/wiki/Statistical_dispersion 20. Gupta, S. P.

(2020). Statistical methods (47th ed.). Sultan Chand & Sons. 21.

Newbold, P., Carlson, W. L., & Thorne, B. (2013).

Statistics for business and economics (8th ed.). Pearson. 22. National Council of Educational Research and Training (NCERT).

Statistics. In Mathematics - Class XI (Chapter 13). NCERT. from https://ncert.nic.in . 23.

OpenStax. (2020). Introduction to statistics .

OpenStax. https://openstax.org/details/books/introductory-statistics 24. Sundar Rao, P. S.

S., & Richard, J. (2018). Correlation and regression analysis.

In Introduction to biostatistics (Revised 4th ed.). PHI Learning Pvt. Ltd. 25.

Daniel, W. W., & Cross, C. L.

(2019). Multiple correlation and regression analysis. In Biostatistics: A foundation for analysis in the health sciences (10th ed.).

Wiley. 26. Mahato, T. K.

(2023). Biostatistical software: Microsoft Excel, SPSS, and Minitab make finding central tendency and dispersion easy and effective. Paripex - Indian Journal of Research, 12 (4). 27.

Mahato, T. K. (2023).

Biostatistics: Simple practical way to use Microsoft Excel to find mean, median, mode, standard deviation, and correlation. World Journal of Biology Pharmacy and Health Sciences, 13 (3), 150–155.

Want the rest of this book?

This preview stops at Chapter 1. Sign up to unlock full chapters, quizzes, and progress tracking.