Data Mining: A Knowledge Discovery Approach Ebook Rar

0 views
Skip to first unread message

Shanae Maerz

unread,
Jul 19, 2024, 10:54:18 PM7/19/24
to dowsverloccligh

In response to this fast-growing demand, universities and colleges have developed data science or data studies majors. These fields have grown from the confluence of statistics, machine learning, AI, and computer science. They are products of a structural transformation in the nature of research in disciplines that include communication, psychology, sociology, political science, economics, business and commerce, environmental science, linguistics, and the humanities. Data mining projects not only require that users possess in-depth knowledge about data processing, database technology, and statistical and computational algorithms; they also require domain-specific knowledge (from experts such as psychologists, economists, sociologists, political scientists, and linguists) to combine with available data mining tools to discover valid and meaningful knowledge. On many university campuses, social sciences programs have joined forces to consolidate course offerings across disciplines to teach introductory, intermediate, and advanced courses on data description, visualization, mining, and modeling to students in the social sciences and humanities.

This chapter examines the major concepts of big data, knowledge discovery in databases, data mining, and computational social science. It analyzes the characteristics of these terms, their central features, components, and research methods.

Data Mining: A Knowledge Discovery Approach ebook rar


DOWNLOAD >>>>> https://blltly.com/2zts1s



There are just as many scholars who think big data is a multifaceted and complex concept that cannot be viewed simply from a data or technology perspective (Mauro, Greco, and Grimaldi 2016). A word cloud analysis from the literature shows that big data can be viewed from at least four different angles. First, big data contains information. The foundation of big data is the production and utilization of information from text, online records, GPS locations, online forums, and so on. This enormous amount of information is digitized, compiled, and stored on computers (Seife 2015). Second, big data includes technology. The enormous size and complexity of the data pose difficulties for computer storage, data processing, and data mining technologies. The technology component of big data includes distributed data storage, cloud computing, data mining, and artificial intelligence. Third, big data encompasses methods. Big data requires a series of processing and analytical methods that are beyond the traditional statistical approaches, such as association, classification, cluster analysis, natural language processing, neural networks, network analysis, pattern recognition, predictive modeling, spatial analysis, statistics, supervised and unsupervised learning, and simulation (Manyika et al. 2011). And fourth, big data has impacts. Big data has affected many dimensions of our society. It has revolutionized how we conduct business, research, design, and production. It has brought and will continue to bring changes in laws, guidelines, and policies on the utility and management of personal information.

Knowledge discovery in database (KDD) is the nontrivial process of identifying valid, novel, potentially useful, and ultimately understandable patterns in data (Fayyad, Piatesky-Shapiro, and Smyth 1996, 84). It consists of nine steps that begin with the development and understanding of the application domain and ends with actions on the knowledge discovered, as illustrated in figure 1.2.

KDD is a nine-step process that involves understanding domain knowledge, selecting a data set, data processing, data reduction, choice of data mining method, data mining, interpreting patterns, and consolidating discovered knowledge. This process is not a one-way flow. Rather, during each step, researchers can backtrack to any of the previous steps and start again. For example, while considering a variety of data mining methods, researchers may go back to the literature and study existing work on the topic to decide which data mining strategy is the most effective one to address the research question.

Data mining has two definitions. The narrow definition is that it is a step in the KDD process of applying data analysis and discovery algorithms to produce certain patterns or models on the data. As shown in figure 1.2, data mining is step 7 in the nine steps of the KDD model. It is usually the case that data mining operates in a pattern space that is infinite, and data mining searches that space to find patterns.

AI, another foundation of data mining techniques, contributes to data mining development with information processing techniques based on a human reasoning model that is heuristic. Machine learning represents an important approach in data mining that trains computers to recognize patterns in data. An artificial neural network (ANN) consists of structures of a large number of highly interconnected processing elements (neurons) working in unison to solve specific problems. ANNs are modeled after human brains in how they process information, and they learn by example by adjusting to the synaptic connections that exist between the neurons.

This book uses the terms knowledge discovery and data mining interchangeably, according to the broadest conceptualization of data mining. Knowledge discovery and data mining in the social sciences constitute a research process that is guided by social science theories. Social scientists with deep domain knowledge work alongside data miners to select appropriate data, process the data, and choose suitable data mining technologies to conduct visualization, analysis, and mining of data to discover valid, novel, potentially useful, and ultimately understandable patterns. These new patterns are then consolidated with existing theories to develop new knowledge. Knowledge discovery and data mining in the social sciences are also important components of computational social science.

Computational social science (CSS) is a new interdisciplinary area of research at the confluence of information technology, big data, social computing, and the social sciences. The concept of computational social science first gained recognition in 2009 when Lazer and colleagues (2009) published Computational Social Science in the journal Science. Email, mobile devices, credit cards, online invoices, medical records, and social media have recorded an enormous amount of long-term, interactive, and large-scale data on human interactions. CSS is based on the collection and analysis of big data and the use of digitalization tools and methods such as social computing, social modeling, social simulation, network analysis, online experiments, artificial intelligence to research human behaviors, collective interactions, and complex organizations (Watts 2013). Only computational social science can provide us with the unprecedented ability to analyze the breadth and depth of vast amounts of data, thus affording us a new approach to understanding individual behaviors, group interactions, social structures, and societal transformations.

Another conceptualization postulates that CSS has four important characteristics. First, it uses data from natural samples that document actual human behaviors (unlike the more artificial data collected from experiments and surveys). Second, the data are big and complex. Third, patterns of individual behavior and social structure are extracted using complex computations based on cloud computing with big databases and data mining approaches. And fourth, scientists use theoretical ideas to guide data mining of big data (Shah et al. 2015).

Others believe that CSS should be an interdisciplinary area at the confluence of domain knowledge, data management, data analysis, and transdisciplinary collaboration and coordination among scholars from different disciplinary training (Mason, Vaughan, and Wallach 2014). Social scientists provide insights on research background and questions, deciding on data sources and methods of collection, while statisticians and computer scientists develop appropriate mathematical models and data mining methods, as well as the necessary computational knowledge and skills to maintain smooth project progress.

Methods of computational social science consist primarily of social computing, online experiments, and computer simulations (Conte 2016). Social computing uses information processing technology and computational methods to conduct data mining and analysis on big data to reveal hidden patterns of collective and individual behaviors. Online experiments as a new research method use the internet as a laboratory to break free of the confines of conventional experimental approaches and use the online world as a natural setting for experiments that transcend time and space (Bond et al. 2012; Kramer, Guillory, and Hancock 2014). Computer simulations use mathematical modeling and simulation software to set and adjust program parameters to simulate social phenomena and detect patterns of social behaviors (Bankes 2002; Gilbert et al. 2005; Epstein 2006). Both online experiments and computer simulations emphasize theory testing and development.

As figure 1.5 shows, in CSS researchers operate under the guidance of social scientific theories, apply computational social science methodology to data (usually big data) from natural samples, detect hidden patterns to enrich social science empirical evidence, and contribute to theory discovery.

The book has six parts. Part I, comprising this chapter and chapter 2, explains the concepts and development of data mining and knowledge and the role it plays in social science research. Chapter 2 provides information on the process of scientific research as theory-driven confirmatory hypothesis testing. It also explains the impact of the new data mining and knowledge discovery approaches on this process.

Part III focuses on model assessment. Chapter 5 explains important methods and measures of model selection and model assessment, such as cross-validation and bootstrapping. It provides justifications as well as ways to use these methods to evaluate models. This chapter is more challenging than the previous chapters. Because the content is difficult for the average undergraduate student, I recommend that instructors selectively introduce sections of this chapter to their students. Later chapters on specific approaches also introduce some of these model assessment approaches. It may be most effective to introduce these specific methods of model assessment after students acquire knowledge of these data mining techniques.

Reply all
Reply to author
Forward
0 new messages