Math & Data Science
11 questions · Fundamental Engineering
From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array indexes start at 1.
Percentile is a statistical measure used to describe the position of a particular value in a dataset relative to the remainder of the values. In education, percentiles are commonly used to rank students’ performance. For instance, if a student is in the 90th percentile, it implies the score of the student is higher than 90% of the other students. To calculate percentile of a student, use the following formula:
For instance, if the scores of five students are given as an array {30, 60, 80, 72, 15}, then the percentile array for each student will be {25, 50, 100, 75, 0}. This implies that the number of students with scores less than 30 is 1; thus, the percentile for the first student is (1 ÷ (5 - 1) × 100) = 25, the number of students with scores less than 60 is 2, and the percentile for the second student is (2 ÷ (5 - 1) × 100) = 50. The same applies thereafter.
The function calculatePercentile receives the scores of each student in the argument arr and calculates and returns the percentile of each student. Assume that the argument arr contains two or more numbers between 0 and 100.
Answer group
From the answer group below, select the correct combination of answers to be inserted into A and B in the description.
Bag of Words is a representation of text that describes the occurrence of words in a document. It is a feature extraction method in text mining. From a document collection (or corpus), a feature vector for each document can be created by considering the count of each word in that document.
The following are two documents that construct a text corpus:
document-I: John likes to watch movies. Mary likes movies too.
document-II: Mary also likes to watch football games.
Here, the word sequence, defined by the unique words (ignoring case and punctuation) in the documents, is as follows:
{John, Mary, likes, to, watch, movies, too, also, football, games}
Now, feature vectors for document-I and document-II can be created using the unique words.
Feature Vector for document-I = {1, 1, 2, 1, 1, 2, 1, 0, 0, 0}
Feature Vector for document-II = {0, 1, 1, 1, 1, 0, 0, 1, 1, 1}
Again, two new text documents and the corresponding word sequence are given below:
document-III: He is a good boy. An honest boy is good.
document-IV: He is an honest boy. The good boy is honest.
Here, the word sequence for those two documents is {He, is, a, an, good, boy, honest, the}. Then, the Feature Vectors for document-III and document-IV will be A and B, respectively.
Answer group
From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array index starts at 1.
Inverse document frequency (IDF) is a metric used to evaluate the importance of a word in a document relative to a document collection. The function calcIDF receives a string term and an array of strings corpus as arguments, and calculates and returns the IDF score for the term based on the given corpus. The corpus of documents corpus is given as a string array. An example of corpus that indicates 3 documents is as follows:
The table shows the description of the functions used in the program.
Table Functions
Traditionally, IDF is computed using the following formula:
IDF(t) = log ( N1 + df(t) )
Where N is the total number of documents, t refers to a specific term or word in a document, and df(t) is the number of documents containing the term t.
For example, the return value of calcIDF( is log(3 ÷ (1 + 2)) = 0 when array corpus is assigned as shown in the above example. Because the total number of documents is 3, and the total number of documents that include the word “ITPEC” is 2.
Answer group
From the answer group below, select the correct combination of answers to be inserted into blank in the program. Here, the array index starts at 0.
The Jaccard similarity index, J, measures the similarity between two sets by computing the ratio of the number of elements in their intersection to the number of elements in their union, i.e., J(A, B) = |A∩B| / |A∪B|, where A and B represent sets, ∩ and ∪ represent the intersection and the union set operations, respectively, and |…| is the operator to calculate the number of elements in a set. The intersection operation builds a new set taking members that are common to both sets; while the union operation builds a set taking all members of both sets, where each member will appear only once. For instance, A = {3, 4, 6, 7} and B = {2, 4, 5}, then A∩B = {4}, A ∪ B = {2, 3, 4, 5, 6, 7}, |A| = 4, | B| = 3, |A∩B| = 1, and |A∪B| = 6.
The function jaccardSimilarity calculates the Jaccard similarity stated above and returns it. The definitions of variables used in the function are stated in the below table.
Table Variable definitions
Answer group
From the answer group below, select the correct combination of answers to be inserted into A and B in the description. Here, the array index starts at 1.
The calcDistance returns a value based on array p1, p2 and positive integer n. Assuming that the number of elements in arrays p1 and p2 are the same. The function abs( returns the absolute value of x, and pow( returns a raised to the power of b.
In the case of p1 is {3, 1, 5, 2} and p2 is {4, 6, 2, 3}, the calcDistance( returns A, whereas calcDistance( returns B. As the value of n increases, those of calcDistance( converges to 5.
Answer group
From the answer group below, select the correct combination of answers to be inserted into A through C in the program.
Craps is a casino dice game in which players bet on the outcomes of the roll of a pair dice. In the rules of the game, a player rolls two dice and finds their sum. There possible outcomes could lead to a win or loss.
The following program approximates the chance of winning a game of Craps using Monte Carlo method and outputs it. The Monte Carlo method is a technique used to estimate the probability of certain outcomes of an experiment by running multiple trial runs, using random numbers. In the program, the playing Craps is illustrated simply by generating random numbers rather than actually rolling a pair of dice. The function random_int( generates a random integer number between 1 and 6, and returns it.
The program also calculates and outputs the relative error of the measured approximate probability of winning Craps. In general, the relative error Er is calculated using the following formula:
Er = | (Pm – Pt) / Pt |
where Pm denotes the measured approximate probability and Pt denotes the theoretical probability. The theoretical probability of winning Craps is known to be 244/495.
Answer group
From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array indexes start at 1.
Bicycle ridership in a city is studied. Examination of several years of data revealed that 30% of the people who regularly ride bicycles in a given year do not regularly ride bicycles in the subsequent year. Additionally, 2% of the people who do not regularly ride bicycles in that year begin to ride bicycles regularly in the subsequent year. If 5,000 people ride bicycles and 100,000 people do not ride bicycles in a given year, then the following program calculates the number of cyclists in the subsequent year (cycling) and that of people who do not ride bicycles in the subsequent year (noncycling):
Answer group
From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array index starts at 1.
Function calcGeoMean receives the array dataArray (the number of elements ≥ 1) as an argument and returns the geometric mean of values in dataArray as the return value. For the numbers a1, a2, …, an, the geometric mean is n√(a1 × a2 × ⋯ × an). Here, the pow( function returns a real value by raising x to the power of y. In the program, division is performed in data type real.
Answer group
From the answer group below, select the correct answer to be inserted into blank in the description. Here, the array indexes start at 1.
The function summarize receives the array sortedData sorted in ascending order and returns five values that characterize the array. The array sortedData must have at least one element. The summarize calls the findRank with two arguments, sortedData and q.
When the function summarize is called as summarize(, the return value is blank.
Answer group
From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array indexes start at 1.
Suppose the weather in a city, from day to day, follows a Markov process. This implies the next day’s weather depends only on today’s weather and not on those of the previous days. If it rains today in the city, a probability 0.3 exists that it will rain again tomorrow. Similarly, the probability that tomorrow’s weather will be sunny is 0.4 and that it will be cloudy is 0.3. If today’s weather is sunny, the next day will be rainy, sunny, or cloudy with probabilities 0.2, 0.7, and 0.1, respectively. Additionally, if today’s weather is cloudy, the next day will be rainy, sunny, or cloudy with probabilities 0.25, 0.5, and 0.25, respectively.
Let us say that the weather is in state 1 if it is rainy, in state 2 if it is sunny, and in state 3 if it is cloudy. Then, the above is a three-state Markov chain whose transition probabilities are given by the following:
Using this transition matrix, the probability distribution of the weather after several days can be estimated.
The program calculates the probability distribution of the weather of the city after three days from the present day and outputs a 3 × 3 square matrix that shows this probability distribution of the weather. For instance, if the weather on the present day is sunny (state 2), the probability that the weather three days after will be rainy (state 1) is the value in the second row and the first column of the outputted matrix.
Answer group
From the answer group below, select the correct answer to be inserted into blank in the description. Here, the array indexes start at 1.
The function calcSim takes two vectors (arrays) as input, calculates their similarity, and returns a value that characterizes the similarity. A large similarity indicates that the vectors are similar, and a small similarity indicates that the vectors are dissimilar. For example, to quantify the similarity between two documents, you can create a vector of occurrences of some words for each of the two documents and input them into calcSim. When the function calcSim( is called, the return value rounded to the first decimal place is blank.
Answer group