Math & Data Science

Math & Data Science

11 questions · Fundamental Engineering

Practice

From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array indexes start at 1.

Percentile is a statistical measure used to describe the position of a particular value in a dataset relative to the remainder of the values. In education, percentiles are commonly used to rank students’ performance. For instance, if a student is in the 90th percentile, it implies the score of the student is higher than 90% of the other students. To calculate percentile of a student, use the following formula:

percentile = number of students who scored below the target student’s scoretotal number of students − 1 × 100

For instance, if the scores of five students are given as an array {30, 60, 80, 72, 15}, then the percentile array for each student will be {25, 50, 100, 75, 0}. This implies that the number of students with scores less than 30 is 1; thus, the percentile for the first student is (1 ÷ (5 - 1) × 100) = 25, the number of students with scores less than 60 is 2, and the percentile for the second student is (2 ÷ (5 - 1) × 100) = 50. The same applies thereafter.

The function calculatePercentile receives the scores of each student in the argument arr and calculates and returns the percentile of each student. Assume that the argument arr contains two or more numbers between 0 and 100.

[Program]
○ real []: calculatePercentile(integer []: arr)
integer: n ← the number of elements of arr
real []: percentiles ← {n zeros}
integer: i, j, count
for (increase i from 1 to n by 1)
count ← 0
for (increase j from 1 to A by 1)
if (arr[i] > arr[j])
count ← count + 1
endif
endfor
percentiles[i] ← B
endfor
return percentiles

Answer group

OptionAB

From the answer group below, select the correct combination of answers to be inserted into A and B in the description.

Bag of Words is a representation of text that describes the occurrence of words in a document. It is a feature extraction method in text mining. From a document collection (or corpus), a feature vector for each document can be created by considering the count of each word in that document.

The following are two documents that construct a text corpus:

document-I: John likes to watch movies. Mary likes movies too.

document-II: Mary also likes to watch football games.

Here, the word sequence, defined by the unique words (ignoring case and punctuation) in the documents, is as follows:

{John, Mary, likes, to, watch, movies, too, also, football, games}

Now, feature vectors for document-I and document-II can be created using the unique words.

Feature Vector for document-I = {1, 1, 2, 1, 1, 2, 1, 0, 0, 0}

Feature Vector for document-II = {0, 1, 1, 1, 1, 0, 0, 1, 1, 1}

Again, two new text documents and the corresponding word sequence are given below:

document-III: He is a good boy. An honest boy is good.

document-IV: He is an honest boy. The good boy is honest.

Here, the word sequence for those two documents is {He, is, a, an, good, boy, honest, the}. Then, the Feature Vectors for document-III and document-IV will be A and B, respectively.

Answer group

OptionAB

From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array index starts at 1.

Inverse document frequency (IDF) is a metric used to evaluate the importance of a word in a document relative to a document collection. The function calcIDF receives a string term and an array of strings corpus as arguments, and calculates and returns the IDF score for the term based on the given corpus. The corpus of documents corpus is given as a string array. An example of corpus that indicates 3 documents is as follows:

{"ITPEC includes members from 6 countries",
"many students prepare for the exam",
"The ITPEC exam is essential for IT professionals"}

The table shows the description of the functions used in the program.

Table Functions

FunctionReturn valueDescription
split(string: str)string []Returns words separated by a space in the text str.
log(real: x)realReturns natural logarithm (base e) of the value x.

Traditionally, IDF is computed using the following formula:

IDF(t) = log ( N1 + df(t) )

Where N is the total number of documents, t refers to a specific term or word in a document, and df(t) is the number of documents containing the term t.

For example, the return value of calcIDF("ITPEC", corpus) is log(3 ÷ (1 + 2)) = 0 when array corpus is assigned as shown in the above example. Because the total number of documents is 3, and the total number of documents that include the word “ITPEC” is 2.

[Program]
○ real: calcIDF(string: term, string []: corpus)
integer: i, j, numDocs, numWords, termCount
boolean: isContainsTerm
real: idf
string []: words
termCount ← 0
numDocs ← the number of elements in corpus
for (increase i from 1 to numDocs by 1)
isContainsTerm ← false
words ← split(corpus[i])
numWords ← the number of elements in words
for (increase j from 1 to numWords by 1) // α
if (A)
isContainsTerm ← true
exit the for block marked α
endif
endfor
if (B)
termCount ← termCount + 1
endif
endfor
idf ← log(numDocs ÷ (1 + termCount)) /* Division is performed
in data type real */
return idf

Answer group

OptionAB

From the answer group below, select the correct combination of answers to be inserted into blank in the program. Here, the array index starts at 0.

The Jaccard similarity index, J, measures the similarity between two sets by computing the ratio of the number of elements in their intersection to the number of elements in their union, i.e., J(A, B) = |A∩B| / |A∪B|, where A and B represent sets, ∩ and ∪ represent the intersection and the union set operations, respectively, and |…| is the operator to calculate the number of elements in a set. The intersection operation builds a new set taking members that are common to both sets; while the union operation builds a set taking all members of both sets, where each member will appear only once. For instance, A = {3, 4, 6, 7} and B = {2, 4, 5}, then A∩B = {4}, A ∪ B = {2, 3, 4, 5, 6, 7}, |A| = 4, | B| = 3, |A∩B| = 1, and |A∪B| = 6.

The function jaccardSimilarity calculates the Jaccard similarity stated above and returns it. The definitions of variables used in the function are stated in the below table.

Table Variable definitions

VariableExplanation
nANumber of elements in set A.
nBNumber of elements in set B.
iCountVariable for counting members of intersection set, i.e., for counting |A∩B|.
uCountVariable for counting members of union set, i.e., for counting |A∪B|.
foundVariable used to indicate whether the member has already been considered for union set or not.
[Program]
○ real: jaccardSimilarity(integer []: A, integer []: B)
integer: nA ← number of elements in A
integer: nB ← number of elements in B
integer: iCount ← 0
integer: uCount ← blank
boolean: found
for (increase i from 0 to nA - 1 by 1)
found ← false
for (increase j from 0 to nB - 1 by 1) // α
if (A[i] = B[j])
iCount ← iCount + 1
found ← true
exit the for block marked α
endif
endfor
if (found = false)
uCount ← uCount + 1
endif
endfor
return iCount ÷ uCount /* Division is done in data type real */

Answer group

From the answer group below, select the correct combination of answers to be inserted into A and B in the description. Here, the array index starts at 1.

The calcDistance returns a value based on array p1, p2 and positive integer n. Assuming that the number of elements in arrays p1 and p2 are the same. The function abs(x) returns the absolute value of x, and pow(a, b) returns a raised to the power of b.
In the case of p1 is {3, 1, 5, 2} and p2 is {4, 6, 2, 3}, the calcDistance(p1, p2, 1) returns A, whereas calcDistance(p1, p2, 2) returns B. As the value of n increases, those of calcDistance(p1, p2, n) converges to 5.

[Program]
○ real: calcDistance(real []: p1, real []: p2, integer: n)
integer: i
real: distance ← 0
real: ex
for (increase i from 1 to number of elements in p1 by 1)
distance ← distance + pow(abs(p1[i] - p2[i]), n)
endfor
ex ← 1 ÷ n /* Division is performed in data type real */
distance ← pow(distance, ex)
return distance

Answer group

OptionAB

From the answer group below, select the correct combination of answers to be inserted into A through C in the program.

Craps is a casino dice game in which players bet on the outcomes of the roll of a pair dice. In the rules of the game, a player rolls two dice and finds their sum. There possible outcomes could lead to a win or loss.

1.If the sum is 7 or 11, the player wins.2.If the sum is 2, 3 or 12, the player loses.3.Otherwise (the sum is 4, 5, 6, 8, 9 or 10), the player neither wins nor loses. However, the player continues rolling the dice until they either roll the same (initial) sum again (in which case they win the game) or roll a sum of 7 (in which case they lose the game).

The following program approximates the chance of winning a game of Craps using Monte Carlo method and outputs it. The Monte Carlo method is a technique used to estimate the probability of certain outcomes of an experiment by running multiple trial runs, using random numbers. In the program, the playing Craps is illustrated simply by generating random numbers rather than actually rolling a pair of dice. The function random_int(1, 6) generates a random integer number between 1 and 6, and returns it.
The program also calculates and outputs the relative error of the measured approximate probability of winning Craps. In general, the relative error Er is calculated using the following formula:

Er = | (Pm – Pt) / Pt |

where Pm denotes the measured approximate probability and Pt denotes the theoretical probability. The theoretical probability of winning Craps is known to be 244/495.

[Program]
integer: wins_sum ← 0
integer: lose_sum ← 0
integer: n ← 10000
integer: i, dice1, dice2, sum, newsum
real: result, pt ← (244 ÷ 495)
for (increase i from 1 to n by 1)
dice1 ← random_int(1, 6)
dice2 ← random_int(1, 6)
sum ← dice1 + dice2
if (sum = 7 or sum = 11)
wins_sum ← wins_sum + 1
elseif (sum = 2 or sum = 3 or sum = 12)
lose_sum ← lose_sum + 1
else
do
dice1 ← random_int(1, 6)
dice2 ← random_int(1, 6)
newsum ← dice1 + dice2
if (newsum = sum)
wins_sum ← wins_sum + 1
elseif (newsum = 7)
lose_sum ← lose_sum + 1
endif
while (A)
endif
endfor
result ← B
output result, absolute value of (C)

Answer group

OptionABC

From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array indexes start at 1.

Bicycle ridership in a city is studied. Examination of several years of data revealed that 30% of the people who regularly ride bicycles in a given year do not regularly ride bicycles in the subsequent year. Additionally, 2% of the people who do not regularly ride bicycles in that year begin to ride bicycles regularly in the subsequent year. If 5,000 people ride bicycles and 100,000 people do not ride bicycles in a given year, then the following program calculates the number of cyclists in the subsequent year (cycling) and that of people who do not ride bicycles in the subsequent year (noncycling):

[Program]
real: pc ← 0.3
real: pb ← 0.02
real: pa ← (1 - pc)
real: pd ← (1 - pb)
integer []: N ← {5000, 100000}
real [,]: P ← {{pa, pb}, {pc, pd}}
real: cycling, noncycling
cycling ← A
noncycling ← B
output cycling, noncycling

Answer group

OptionAB

From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array index starts at 1.

Function calcGeoMean receives the array dataArray (the number of elements ≥ 1) as an argument and returns the geometric mean of values in dataArray as the return value. For the numbers a1, a2, …, an, the geometric mean is n√(a1 × a2 × ⋯ × an). Here, the pow(x, y) function returns a real value by raising x to the power of y. In the program, division is performed in data type real.

[Program]
○ real: calcGeoMean(real []: dataArray)
real: product, geomean
integer: n ← the number of elements in dataArray
integer: i
product ← 1
for (increase i from 1 to n by 1)
product ← product × dataArray[i]
endfor
geomean ← pow(A, B)
return geomean

Answer group

OptionAB

From the answer group below, select the correct answer to be inserted into blank in the description. Here, the array indexes start at 1.

The function summarize receives the array sortedData sorted in ascending order and returns five values that characterize the array. The array sortedData must have at least one element. The summarize calls the findRank with two arguments, sortedData and q.
When the function summarize is called as summarize({0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1}), the return value is blank.

[Program]
○ real: findRank(real []: sortedData, real: q)
integer: j
j ← floor(q × (the number of elements in sortedData - 1))
// floor returns the closest integer less than or equal to a given
// number, e.g. floor(7.75) returns 7.
return sortedData[j + 1]
○ real []: summarize(real []: sortedData)
real []: rankData ← {} /* array with 0 elements */
real []: q ← {0, 0.25, 0.5, 0.75, 1}
integer: i
for (increase i from 1 to the number of elements in q by 1)
add the return value of findRank(sortedData, q[i]) to the end of rankData
endfor
return rankData

Answer group

From the answer group below, select the correct combination of answers to be inserted into A and B in the program. Here, the array indexes start at 1.

Suppose the weather in a city, from day to day, follows a Markov process. This implies the next day’s weather depends only on today’s weather and not on those of the previous days. If it rains today in the city, a probability 0.3 exists that it will rain again tomorrow. Similarly, the probability that tomorrow’s weather will be sunny is 0.4 and that it will be cloudy is 0.3. If today’s weather is sunny, the next day will be rainy, sunny, or cloudy with probabilities 0.2, 0.7, and 0.1, respectively. Additionally, if today’s weather is cloudy, the next day will be rainy, sunny, or cloudy with probabilities 0.25, 0.5, and 0.25, respectively.
Let us say that the weather is in state 1 if it is rainy, in state 2 if it is sunny, and in state 3 if it is cloudy. Then, the above is a three-state Markov chain whose transition probabilities are given by the following:

to123from10.30.40.320.20.70.130.250.50.25

Using this transition matrix, the probability distribution of the weather after several days can be estimated.
The program calculates the probability distribution of the weather of the city after three days from the present day and outputs a 3 × 3 square matrix that shows this probability distribution of the weather. For instance, if the weather on the present day is sunny (state 2), the probability that the weather three days after will be rainy (state 1) is the value in the second row and the first column of the outputted matrix.

[Program]
real [,]: pmat ← {{0.3, 0.4, 0.3}, {0.2, 0.7, 0.1}, {0.25, 0.5, 0.25}}
real [,]: dmat ← pmat
real [,]: smat
integer: i, j, k, m
for (increase m from 1 to A by 1)
smat ← {{0, 0, 0}, {0, 0, 0}, {0, 0, 0}}
for (increase i from 1 to 3 by 1)
for (increase j from 1 to 3 by 1)
for (increase k from 1 to 3 by 1)
smat[i, j] ← smat[i, j] + B
endfor
endfor
endfor
dmat ← smat
endfor
output dmat

Answer group

OptionAB

From the answer group below, select the correct answer to be inserted into blank in the description. Here, the array indexes start at 1.

The function calcSim takes two vectors (arrays) as input, calculates their similarity, and returns a value that characterizes the similarity. A large similarity indicates that the vectors are similar, and a small similarity indicates that the vectors are dissimilar. For example, to quantify the similarity between two documents, you can create a vector of occurrences of some words for each of the two documents and input them into calcSim. When the function calcSim({2, 2, 1, 0, 4}, {3, 1, 1, 1, 2}) is called, the return value rounded to the first decimal place is blank.

[Program]
// Assume that arrays v1 and v2 have the same number of one or more elements
// and that the arrays are not all-zero.
○ real: calcSim(integer []: v1, integer []: v2)
integer: i, x, y
integer: sxx ← 0
integer: syy ← 0
integer: sxy ← 0
for (increase i from 1 to the number of elements in v1 by 1)
x ← v1[i]
y ← v2[i]
sxx ← sxx + x × x
syy ⟵ syy + y × y
sxy ⟵ sxy + x × y
endfor
return sxy ÷ (square root of (sxx × syy))

Answer group