IS5113 All Week’s quizzes

  1. One of the key differences between business analytics and data science is their primary focus either on business problems or on mathematical algorithms.
    A) True
  2. Analytics and analysis are essentially the same thing; they both focus on the granular level representation of complex problems through decomposition of the whole into its lower-level parts.
    A) False
  3. If a data scientist is analyzing historical data to identify problems and root causes, he/she is essentially conducting descriptive analytics.
    A) True
  4. ERP stands for enterprise resource planning and is used for the integration of company-wide data.
    A) True
  5. The most important driver behind business analytics popularity is the need for business managers to make experience and intuition driven business decisions.
    A) False
  6. Business analytics and data science have the same purpose: to convert data into actionable insight through an algorithm-based discovery process.
    A) True
  7. Major commercial business intelligence products and services were established in the early 1970s.
    A) False
  8. If I am distributing funds to different financial products to maximize return, I am essentially doing descriptive analytics.
    A) False
  9. Today, analytics can be defined simply as “the discovery of information/knowledge/insight in data.”
    A) True
  10. Business intelligence is a broad concept that also includes business analytics within its simple taxonomy.
    A) False
  11. Analytics is the art and science of discovering insight to support accurate and timely decision making.
    A) True
  12. Business analytics is the process of developing computer code and novel IT frameworks.
    A) False
  13. Organizations apply analytics to business problems to identify problems, foresee future trends, and make the best possible decisions.
    A) True
  14. DeepQA is a massively parallel, web mining focused, probabilistic computational algorithm developed by the SAS Institute
    A) False
  15. Descriptive analytics is also called business intelligence that is the entry level in analytics taxonomy.
    A) True
  16. What are the main roadblocks to the adoption of analytics?
    A) All of these
  17. Jim, the marketing manager in the company, is interested in the sales numbers in the south region by each product type for the last six months. What type of analytics would you use to help him?
    A) Descriptive
  18. Which of the following developments is not contributing to facilitating the growth of decision support and analytics? AO Knowledge management systems
    A) Locally concentrated workforces
  19. What type of analytics seeks to identify the courses of action to achieve the best performance possible?
    A) Prescriptive
  20. If Jack is interested in identifying the optimal quantity of purchase orders in order to minimize the overall cost, which of the following types of analytics should he use?
    A) Prescriptive
  21. Firms have used analytics to enhance which of the following business activities?
    A) All of these
  22. Which of the following is not commonly used as an enabler of descriptive analytics?
    A) Data mining
  23. Association patterns can include capturing the sequence of events and things.
    A) True
  24. Cubes in OLAP are defined as a multidimensional representation of the data stored in and retrieved from data warehouses.
    A) True
  25. Prediction modeling is often classified under the unsupervised machine learning methods.
    A) False
  26. Data mining can be used to predict the result of sporting events to identify means to decrease odds of winning against specific opponent.
    A) False
  27. In banking and finance, data mining is often used to manage microeconomics movements and overall cash flow outcomes.
    A) False
  28. One of the most pronounced reasons for the increasing popularity of data mining is due to the fact that there are less suppliers than corresponding demand in the business marketplace.
    A) False
  29. Novel is a key term in the definition of data mining, which means that the patterns are known by the user within the context of the system being analyzed.
    A) False
  30. Segmentation and outlier analysis are part of classification modeling.
    A) False
  31. Data mining is primarily concerned with mining (that is, digging out data) from a variety of disparate data sources.
    A) False
  32. In the retail industry, association rule mining is frequently called market-based analysis.
    A) True
  33. CRM aims to create one-on-one relationships with customers by developing an intimate understanding of their needs and wants.
    A) True
  34. Data mining leverages capabilities of statistics, artificial intelligence, machine learning, management science, information systems, and databases in a systematic and synergistic way.
    A) True
  35. The original terminology of data mining commonly refers to discovering known patterns in large and structured data sets.
    A) False
  36. Manufacturers use data mining to classify anomalies and commonalities in the production system to improve the manufacturing system.
    A) True
  37.  Information warfare often refers to identify and stop malicious attacks on critical information infrastructures in literarily any and every organizations and business
    A) True
  38. In data mining, clustering is classified further into:
    A) segmentation and outlier analysis.
  39. Which of the following is the most commonly used clustering
    A) k-means
  40. What kinds of patterns can data mining discover?
    Each correct answer represents a complete solution. Choose all that apply.
    A) Clustering
    Classification
    Forecasting
    Association
  41. What are the most common reasons why data mining has gained overwhelming attention in the business world?
    A) All of these
  42. In retailing, data mining is most commonly used to:
    A)predict future sales.
  43. predict future sales.
    A) Assigns customers to different segments
  44. What is the primary difference between statistics and data mining?
    A) Statistics starts with a well-defined proposition and hypothesis, whereas data mining starts with a loosely defined discovery statement.
  45. The important part of the KDD process is the feedback loop that allows the process flow to redirect backward, from any step to any other previous steps, for rework and readjustments.
    A) True
  46. The data sources that are combined in a centralized data repository for supporting managerial decisions is known as a data warehouse.
    A) True
  47. In the SEMMA process, the accuracy and usefulness of the models are evaluated in the Assess step.
    A) True
  48. In the SEMMA process, visualization and description of the data are carried out in the Modify step.
    A) False
  49. The CRISP-DM methodology was proposed by Fayyad et al., in the year 1996.
    A) False
  50. In the model building task, both the CRISP-DM and SEMMA methodologies build and test various models.
    A) True
  51. Define, Explore, Measure, and Assess are the steps involved in the Six Sigma process.
    A) False
  52. In the testing and evaluation step of the CRISP-DM methodology, monitoring and maintenance of the models are important.
    A) False
  53. During the model building step in the CRISP-DM process, the data mining methods and algorithms are applied to the current data set.
    A) True
  54. The Six Sigma process promotes an error-free/perfect business execution.
    A) True
  55. The Modify step in Six Sigma involves the process of assessing the mapping between organizational data repositories and the business problem.
    A) False
  56. In the project finalization task, both the CRISP-DM and SEMMA methodologies prescribe deploying the results.
    A) False
  57. Identifying the most pressing problem and defining the goals and objectives can be done in the Define step of the Six Sigma process.
    A)True
  58. When compared with all other methodologies, CRISP-DM is the most popular data mining process that is being used in data analytics.
    A) True
  59. In the CRISP-DM process, it is not important or necessary to follow the sequential order of each step. That is, the steps can be executed in an arbitrary sequence.
    A) False
  60. During which step of the SEMMA process the analyst searches for unanticipated trends and anomalies to gain a better understanding of the data set?
    A) Explore
  61. Which of the following steps of the CRISP-DM process is commonly called the data preprocessing step that produces the data identified in the data understanding | step for analysis?
    A) Data preparation
  62. Which of the following is the most relevant methodology that is used to implement data science and business analytics projects?
    A) CRISP-DM
  63. During which step of the Six Sigma process are the identified data sources consolidated and transformed into a format that is amenable to machine processing?
    A) Measure
  64. Which of the following steps of the CRISP-DM process identifies the relevant data from different sources?
    A) Data understanding
  65. Which of the following substeps are involved in the Sample step of the SEMMA process?
    A) Training, validation, and test
  66. Which of the following steps of the CRISP-DM process identifies the goals, purpose, and requirements of the customers?
    A) Business understanding
  67. The customer credit ratings like bad, fair, and excellent are considered as what type of data?
    A) Ordinal
  68. The ratio of accurately classified instances (positives and negatives) divided by the total number of instances is defined as the overall accuracy metric.
    A) True
  69. Handling the missing values in the data is typically performed in the data consolidation phase.
    A) False 
  70. F1 metric is simply the harmonic mean of precision and recall.
    A) True
  71. A typical example of interval scale measurement is the temperature on the Celsius scale.
    A) True
  72. Apriori and FP-Growth algorithms are part of the association type data mining tasks.
    A) TrueFalse Apriori and FP-Growth algorithms are part of the association type data mining tasks.
    A) True
  73. The ratio of correctly classified positives divided by the total positive count is defined as a precision metric.
    A) False 
  74. If a classification problem is not binary, you cannot use a confusion matrix to tabulate prediction outcomes.
    A) False 
  75. k-means algorithm is a part of prediction data mining method.
    A) False
  76. The bootstrapping methodology is similar to the leave-one-out methodology, where it can be used to calculate accuracy by leaving out one sample at each iteration of the estimation process.
    A) False
  77. Balancing skewed data means oversampling the more represented class records and undersampling the less represented class records
    A) False
  78. Decision trees are part of the regression type prediction methods.
    A) False
  79. The multi split methodology partitions data into exactly two mutually exclusive subsets called training set and test set.
    A) False
  80. The purpose of data preparation (commonly called data preprocessing) is to eliminate the possibility of GIGO errors.
    A) True
  81. How and what the model concludes on certain predictions is obtained by the interpretability characteristic of the prediction method.
    A) True
  82. The area under the ROC curve is a graphical assessment technique for binary classification problems, in which sensitivity is plotted on the y-axis and the specificity is plotted on the x-axis.
    A) False
  83. Which clustering method is based on the basic idea that nearby objects are more related to each other than are those that are farther away from each other?
    A) Hierarchical
  84. Which cross-validation methodology achieves random sampling of a fixed number of instances from the original data with replacement to construct the training data set?
    A) Bootstrapping
  85. Which classification method use(s) conditional probabilities to build classification models?
    A) Bayesian classifiers
  86. Which of the following is defined as the ratio of correctly classified negatives divided by the total negative count?
    A) Specificity
  87. Which of the following factors refers to a model’s ability to make reasonably accurate predictions, given noisy data or data with missing and erroneous values?
    A) Robustness
  88. Which method takes into account the partial membership of class labels to predefined categories while building models for classification problems?
    A)  Rough sets
  89. Time series is a sequence of data points of interest measured and represented at consecutive and regular time intervals.
    A) True
  90. In linear regression, the independence of errors assumption is also known as homoscedasticity.
    A) False
  91. Multicollinearity can be triggered by having two or more perfectly correlated explanatory variables present in the model.
    A) True
  92. In linear regression, hypothesis testing reveals the existence of relationships between explanatory variables.
    A) False 
  93. The Naive Bayes method requires output variables to have numeric values
    A) False
  94.  In prediction, linear regression uses a mathematical equation to identify additive mathematical relationships between explanatory variables and the response variable
    A) True
  95. In the normality of error assumption of linear regression, the response variables’ values are expected to be randomly distributed.
    A) False
  96. In time-series forecasting, an estimator’s mean squared error measures the average absolute error between the estimated and the actual values.
    A) False
  97. Correlation is meant to represent the linear relationships between two nominal input variables.
    A) False
  98. k-NN is a prediction method used not only for classification but also for regression-type prediction problems.
    A) True
  99. To deploy a developed SVM model, the model coefficients can be extracted and integrated directly into the decision support system.
    A) True
  100. Logistic regression is like linear regression where both of them are used to predict a numeric target variable.
    A) False
  101. Linear regression aims to capture the functional relationships between one or more numeric input variables and a categorical output variable.
    A) False
  102. Homoscedasticity states that the response variables must have the same variance in their error, regardless of the explanatory variables’ values.
    A) True
  103. In the SVM model, normalization’s main benefit is to avoid having attributes in greater numeric ranges and dominate those in smaller numeric ranges.
    A) True 
  104. In prediction analytics, variance refers to the error, and bias refers to the consistency in the predictive accuracy of models applied to other data sets.
    A) False
  105. A data set is imbalanced when the distribution of different classes in the input variables are significantly dissimilar.
    A) False
  106. Overfitting is the notion of making the model too specific to the training data to capture not only the signal but also the noise in the data set.
    A) True
  107. Information fusion type model ensembles utilize meta-modeling called super learners.
    A) False
  108. Bias is often defined as the difference between a model’s prediction output and the actual values for a given prediction problem.
    A) True
  109. Model ensembles are known to be more robust against outliers and noise in the data compared to individual models.
    A)True
  110. Bagging type ensembles can be used in both regression and classification type prediction problems.
    A) True
  111. In explainable AI, the LIME and SHAP methods are considered as global interpreters.
    A) False
  112. Sensitivity analysis based on the leave-one-out methodology can be applied to any predictive analytics method because of its model agnostic implementation methodology.
    A) True
  113. A model with low variance is the one that captures both noise and generalized patterns in the data and therefore produces an overfit model.
    A) False
  114. In ensemble modeling, bagging uses the bootstrap sampling of cases to create a collection of decision trees.
    A) True
  115. Model ensembles are much easier and faster to develop than individual models.
    A) False
  116. In ensemble modeling, boosting builds several independent simple trees for the resultant prediction model.
    A) False
  117. Underfitting is mainly characterized on the bias–variance trade-off continuum as low-bias/low-variance outcome.
    A) False
  118. Sensitivity analysis based on input value perturbation is often used in trained feed-forward neural network modeling, where all of the input variables are numeric and standardized.
    A) True
  119. Clustering is a supervised learning process in which objects are assigned to pre-determined number of artificial groups called clusters.
    A) False 
  120. Text-to-speech is a text processing function that can read textual content and detects and corrects syntactic and semantic errors.
    A) False 
  121. In the context of the text mining process, both structured and unstructured data are extracted from the data sources and converted into context-specific knowledge.
    A) True 
  122. SCM and ERP are the first two beneficiaries of the NLP and WordNet.
    A) False 
  123. In marketing applications, text mining can be used to assess and help predict a customer’s propensity to attrite.
    A) True 
  124. Singular value decomposition help reduce the overall structure of the term-document matrix to a lower dimensional space for further pattern/knowledge discovery.
    A) True 
  125. A polygraph is a non-intrusive deception-detection technique commonly used to assess the level of truthfulness in the textual content.
    A) False 
  126. In the first task of the text mining process, the data is structured and preprocessed to achieve hidden patterns and knowledge nuggets.
    A) False 
  127.  In text mining, associations refer to direct relationships between terms or sets of concepts.
    A) True
  128. Automatic summarization is a program that is used to assign documents into a predefined set of categories.
    A) False 
  129. The main aim of NLP is to move away from word counting to a real understanding and processing of natural human language.
    A) True 
  130. Tokenizing refers to the process of breaking sentences into blocks of text that performs a specific linguistic function.
    A) True 
  131. In the context of text mining, lemmatization is a process of syntactically reducing words to their stem/root form.
    A) False 
  132.  In the term-by-document matrix, the columns represent the terms and the rows represent the documents, and the cells represent the variances.
    A)False 
  133. In the context of text mining, structured data is for humans to process, while unstructured data is for computers to process and understand.
    A) False 
  134. In the context of text mining, which of the following is a part of NLP that studies the internal structure of words (that is, the patterns of word formation within a language or across languages)?
    A) Morphology
  135. Which of the following are the most commonly used normalization methods?
    A) Log, binary, and inverse document frequencies
  136. Which of the following are the best options available to manage the TDM matrix size?
    A) Labor-intensive process, eliminate terms, and singular value decomposition
  137. Which of the following are the common challenges that are associated with the implementation of NLP?
    A) All of these
  138. Which of the following is not among the steps involved in sentiment analysis?
    A) Latent Dirichlet allocation
  139. In the knowledge extraction method of the text mining process, ____________ refers to the natural groping, analysis, and navigation of large text collections, such as web pages.
    A) Clustering
  140. Which of the following applications utilize the capabilities of text mining?
    A) Marketing applications
    Security applications
    Biomedical applications
  141. In which of the following categories of knowledge extraction method is the task of text categorization achieved?
    A) Classification
  142. Hadoop is an open-source framework for processing, storing, and analyzing massive amounts of distributed, wide variety of data.
    A) True 
  143. The term velocity in big data analytics refers to how fast digitized data is created and processed.
    A) True 
  144. Big data comes from a variety of sources within an organization, including marketing and sales transaction, inventory records, financial transaction, and human resources and accounting records.
    A)False 
  145. Hadoop is a batch-oriented computing framework, which implies it does not support real-time data processing and analysis.
    A) True 
  146. A stream in a stream analytics is defined as a discrete and aggregated level of data elements.
    A) False 
  147. MapReduce is a contemporary programming language designed to be used by computer programmers.
    A) False 
  148. Among the variety of factors, the key driver for big data analytics is the business needs at any level, including strategic, tactical, or operational.
    A) True 
  149. Grid computing increases efficiency, lowers total cost, and enhances production by processing computational jobs in a shared, centrally managed ordinary pool of computing resources.
    A) True 
  150. HDFS (Hadoop Distributed File System) was invented before Google developed MapReduce. Hence, the early versions of MapReduce relied on HDFS.
    A) False 
  151. The main benefit of Hadoop is that it allows enterprises to process and analyze large volumes of structured and semi-structured data on specialized hardware.
    A) False 
  152. Hadoop is not just about the volume but also processing of diversity of data types.
    A) True
  153. A data scientist’s main objective is to organize and analyze large amounts of data, to solve complex problems, often using software specifically designed for the task.
    A) True 
  154. The term veracity in big data analytics refers to the processing of different types and formats of data, structured and unstructured.
    A) False 
  155. Hadoop is a replacement for a data warehouse which stores and processes large amounts of structured data.
    A) False 
  156. In typical data stream mining applications, the purpose is to predict the class or value of new instances in the data stream, given some knowledge about the class membership or values of previous instances in the data stream.
    A) True 
  157. The main characteristic of deep learning solutions is that they use AI (artificial intelligence) to understand and organize data, predict the intent of a search query, improve the relevancy of results, and automatically tune the relevancy of results over time. xzc xc
    A) False 
  158.  Human–computer interaction is a critical component of cognitive systems that allows users to interact with cognitive machines and define their needs.
    A) True
  159. Deep learning analytics is a term that refers to the computing−branded technology platforms, such as IBM Watson, that specialize in processing and analyzing large, unstructured data sets.
    A) False 
  160. In a typical neural network, the goal of the testing process is to adjust the network weights and biases such that the network output for each set of inputs is adequately close to its corresponding target value.
    A) False 
  161. Connection weights are the key elements of an artificial neural network (ANN). They produce the final value through the summation and transfer function.
    A) False 
  162. AI (artificial intelligence) has the capability to find hidden patterns in a variety of data sources to identify problems and provide potential solutions.
    A) True
  163. Cognitive computing has the capability to simulate human thought processes to assist humans in finding solutions to complex problems.
    A) True
  164. In artificial neural networks, neurons are processing units, also called processing elements, that perform predefined mathematical operations on the numeric values from the input variables or the other neuron outputs to create and push out their own outputs.
    A) True
  165. The term long short-term memory network refers to a network that is used to remember what happened in the past for a long enough time that it can be leveraged in accomplishing the task when needed.
    A) True
  166. Multilayer perceptron type deep networks are also known as feedforward networks because the flow of information that goes through them is always forwarding, and no feedback connections are allowed.
    A) True
  167. In representation learning, the emphasis is on automatically discovering the features to be used for analytics purposes.
    A) True
  168. Delta (or an error) is defined as the difference between the network weights in two consecutive iterations.
    A) False 
  169. The purpose of artificial intelligence is to augment human capability.
    A) False 
  170. The main characteristic of the convolutional networks is having at least one layer involving a convolution weight function instead of general matrix multiplication.
    A) True
  171. Deep learning is an extension of neural networks that deal with more complicated tasks with a higher level of sophistication by employing many layers of connected neurons.
    A) True
  172. What is the primary focus of machine learning?
    A) Developing algorithms that can learn from data
  173. Which of the following is NOT a type of machine learning algorithm mentioned in the video?
    A) Structured learning
  174. Why is data visualization important in machine learning?
    A) It helps to understand patterns and trends in data
  175. What ethical concern was discussed regarding AI in the video?
    A) Bias in algorithms
  176. Which technology is closely related to high-performance computing in the context of AI?
    A) Quantum computing
  177. What role does pattern recognition play in machine learning?
    A) Identifying similarities and anomalies in data
  178. What are some potential applications of AI discussed in the video? Choose all that apply.
    A) All of the above 
  179. Unsupervised learning algorithms require labeled data for training .
    A) False

Leave a Reply

Your email address will not be published. Required fields are marked *