- 15 Sections
- 89 Lessons
- 40 Hours
Expand all sectionsCollapse all sections
- PERSIAPAN2
- 1. GAINING EARLY INSIGHTS FROM TEXTUAL DATA8
- 2.11.1. Exploratory Data Analysis
- 2.21.2. Introducing the Dataset
- 2.31.3. Blueprint: Getting an Overview of the Data with Pandas
- 2.41.4. Blueprint: Building a Simple Text Preprocessing Pipeline
- 2.51.5. Blueprints for Word Frequency Analysis
- 2.61.6. Blueprint: Finding a Keyword-in-Context
- 2.71.7. Blueprint: Analyzing N-Grams
- 2.81.8. Blueprint: Comparing Frequencies Across Time Intervals and Categories
- 2. EXTRACTING TEXTUAL INSIGHTS WITH APIS3
- 3. SCRAPING WEBSITES AND EXTRACTING DATA17
- 4.13.1. Scraping and Data Extraction
- 4.23.2. Introducing the Reuters News Archive
- 4.33.3. URL Generation
- 4.43.4. Blueprint: Downloading and Interpreting robots.txt
- 4.53.5. Blueprint: Finding URLs from sitemap.xml
- 4.63.6. Blueprint: Finding URLs from RSS
- 4.73.7. Downloading Data
- 4.83.8. Blueprint: Downloading HTML Pages with Python
- 4.93.9. Blueprint: Downloading HTML Pages with wget
- 4.103.10. Extracting Semistructured Data
- 4.113.11. Blueprint: Extracting Data with Regular Expressions
- 4.123.12. Blueprint: Using an HTML Parser for Extraction
- 4.133.13. Blueprint: Spidering
- 4.143.14. Density-Based Text Extraction
- 4.153.15. All-in-One Approach
- 4.163.16. Blueprint: Scraping the Reuters Archive with Scrapy
- 4.173.17. Possible Problems with Scraping
- 4. PREPARING TEXTUAL DATA FOR STATISTICS AND MACHINE LEARNING7
- 5. FEATURE ENGINEERING AND SYNTACTIC SIMILARITY5
- 6. TEXT CLASSIFICATION ALGORITHMS6
- 7.16.1. Introducing the Java Development Tools Bug Dataset
- 7.26.2. Blueprint: Building a Text Classification System
- 7.36.3. Final Blueprint for Text Classification
- 7.46.4. Blueprint: Using Cross-Validation to Estimate Realistic Accuracy Metrics
- 7.56.5. Blueprint: Performing Hyperparameter Tuning with Grid Search
- 7.66.6. Blueprint Recap and Conclusion
- 7. HOW TO EXPLAIN A TEXT CLASSIFIER5
- 8.17.1. Blueprint: Determining Classification Confidence Using Prediction Probability
- 8.27.2. Blueprint: Measuring Feature Importance of Predictive Models
- 8.37.3. Blueprint: Using LIME to Explain the Classification Results
- 8.47.4. Blueprint: Using ELI5 to Explain the Classification Results
- 8.57.5. Blueprint: Using Anchor to Explain the Classification Results
- 8. UNSUPERVISED METHODS: TOPIC MODELING AND CLUSTERING9
- 9.18.1. Our Dataset: UN General Debates
- 9.28.2. Nonnegative Matrix Factorization (NMF)
- 9.38.3. Latent Semantic Analysis/Indexing
- 9.48.4. Latent Dirichlet Allocation
- 9.58.5. Blueprint: Using Word Clouds to Display and Compare Topic Models
- 9.68.6. Blueprint: Calculating Topic Distribution of Documents and Time Evolution
- 9.78.7. Using Gensim for Topic Modeling
- 9.88.8. Blueprint: Using Clustering to Uncover the Structure of Text Data
- 9.98.9. Further Ideas
- 9. TEXT SUMMARIZATION5
- 10. EXPLORING SEMANTIC RELATIONSHIPS WITH WORD EMBEDDINGS4
- 11. PERFORMING SENTIMENT ANALYSIS ON TEXT DATA7
- 12.111.1. Sentiment Analysis
- 12.211.2. Introducing the Amazon Customer Reviews Dataset
- 12.311.3. Blueprint: Performing Sentiment Analysis Using Lexicon-Based Approaches
- 12.411.4. Supervised Learning Approaches
- 12.511.5. Blueprint: Vectorizing Text Data and Applying a Supervised Machine Learning Algorithm
- 12.611.6. Pretrained Language Models Using Deep Learning
- 12.711.7. Blueprint: Using the Transfer Learning Technique and a Pretrained Language Model
- 12. BUILDING A KNOWLEDGE GRAPH6
- 13. USING TEXT ANALYTICS IN PRODUCTION5
- 14.113.1. Blueprint: Using Conda to Create Reproducible Python Environments
- 14.213.2. Blueprint: Using Containers to Create Reproducible Environments
- 14.313.3. Blueprint: Creating a REST API for Your Text Analytics Model
- 14.413.4. Blueprint: Deploying and Scaling Your API Using a Cloud Provider
- 14.513.5. Blueprint: Automatically Versioning and Deploying Builds
- PENUTUPAN2
Bundle materi
Next