Registration is open to participants from different fields. Members of the INCT.DD network can apply for a scholarship.
Registration opens this Monday (22 December) for the INCT.DD/Compolítica 2026 Summer School, which runs from 26 January to 13 February. It offers courses in quantitative and qualitative methods for social research. Interested participants from different fields can register. Members of the INCT.DD network may apply for a scholarship for one course of their choice. Compolítica members are also entitled to a discount, which must be requested by email at secretaria@compolitica.org.
There are eleven fully online courses: Introduction to R, Python, Introduction to Statistics, Natural Language Modeling, Introduction to Network Analysis, Introduction to Social Media Data Collection and Storage, Orange, Introduction to Machine Learning, Content Analysis, Narrative Interviews, and Q Methodology.
Network members interested in a scholarship for one course: CLICK HERE.
For general registration through Sympla: CLICK HERE.
The Summer School is held annually by the National Institute of Science and Technology in Digital Democracy (INCT.DD), in partnership with the Brazilian Association of Researchers in Communication and Politics (Compolítica). The project began in 2019 and has trained more than 600 students.
See the full schedule and course descriptions:
Introduction to Python (15 hours)
26–30 January, 10 a.m.–noon
Professor Sivaldo Pereira
This course is intended for humanities researchers with no programming experience. It provides introductory training in data science and programming for research through Python algorithms, with a focus on the Pandas library. Practical activities will use open government datasets. The course will address the social use and reuse of structured data for public transparency, the design of academic research, the production of reports and studies with data visualization, and the foundations of application development.
Introduction to R (15 hours)
26–30 January, 2–4 p.m.
Professor Crysttian Paixão
R is a highly versatile language used for many different activities. The course covers its main commands and the installation and setup of the working environment. It also introduces key R packages, their main commands, and their applications. Text processing uses various techniques to extract knowledge. Some require a programming language, and R offers many resources to support text processing. The course presents some of these resources and applies them to the analysis of different texts to extract knowledge.
Introduction to Statistics (9 hours)
26, 28, and 30 January, 6–8 p.m.
Professor Lineu Cavazani (UFPR)
The course teaches basic concepts in exploratory data analysis. Participants will learn to interpret results critically and perform basic exploratory procedures and analyses for common variable types, using tables, graphs, and summary measures.
Natural Language Modeling (6 hours)
2 and 3 February, 10 a.m.–noon
Tariq Choucair
This course gives a practical introduction to the main techniques of natural language processing (NLP). It covers basic concepts of linguistic representation and traditional methods of automatic text analysis, with a focus on part-of-speech tagging (POS), named entity recognition (NER), and dependency parsing. Classes combine theoretical foundations with applied examples and discuss the use of these techniques in social science and humanities research.
Introduction to Social Media Data Collection and Storage (6 hours)
2 and 3 February, 2–4 p.m.
Professor João Senna (INCT.DD)
The course gives a brief overview of ways to collect and analyze social media data without programming, using tools such as Zeeschuimer, 4CAT, and YouTube Data Tools.
Narrative Interviews (6 hours)
2 and 3 February, 6–8 p.m.
Professor Luciane Belin
This workshop introduces narrative interviews as a technique for qualitative data collection and the narrative method as an analytical strategy. Across two sessions, it presents the method’s theoretical foundations, its differences from other interview types, and the ethical care required in its use. Activities include theoretical presentations, guided exercises, and practice in pairs. Participants will experience each stage, from developing the initial narrative prompt to conducting the interview, including attentive listening, identifying narrative markers, and initial analysis strategies. The workshops also provide space to discuss common challenges, exchange experiences, and reflect on the role of narratives in producing knowledge, especially in research that seeks a detailed understanding of perspectives, life paths, and experiences.
Orange (9 hours)
4–6 February, 2–4 p.m.
Professor Eduardo Grizenti (UFBA)
We face an unprecedented volume of information. Social sciences benefit from the variety and ready availability of data, but also face the challenge of finding methods to collect and analyze such large amounts. The course introduces Orange, a program that uses Python and provides an intuitive way to perform unsupervised analysis of varied text collections through text mining. It covers how to import, prepare, and explore a corpus through word frequency, word clouds, concordances, topic modeling, and proximity.
Introduction to Network Analysis (9 hours)
9–11 February, 2–4 p.m.
Professor João Guilherme Bastos dos Santos (INCT.DD / Democracia em Xeque)
This course presents theoretical principles and empirical methods for using complex network analysis to study online social media platforms, with a particular focus on countering disinformation campaigns. Main topics include algorithmic and voluntary segmentation; clustering methods and critical assessment of algorithmic results; definitions of network centrality and changes in complex networks over time; problems in identifying bots; and network analysis in messaging apps without visibility algorithms. The course uses mixed methods based on free software and the R programming language, combining network, text, and image analysis.
Introduction to Machine Learning (9 hours)
9–11 February, 6–8 p.m.
Professor Alexandre Teles (INCT.DD)
This course covers machine learning and artificial intelligence, focusing on practical applications of large language models (LLMs). Students will learn about synthetic data generation, including structured JSON data and validation with LLMs. It also covers text classification models, including zero-shot and few-shot classification, Sentence Transformers, and SetFit for binary and multiclass classification. Finally, it explores techniques for training LLMs, distinguishes base models from instruction-fine-tuned models, and covers data preparation in ShareGPT format.
Q Methodology (6 hours)
9–11 February, 6–8 p.m.
Professor Larissa Natalino
Q methodology combines quantitative and qualitative procedures to investigate patterns of subjectivity, understood as different viewpoints on the same topic. Although still little used in communication studies, it has strong potential for analyzing discourse, opinions, and public controversies. Organized in two three-hour sessions, the course presents its theoretical and methodological foundations and discusses applications in academic research. It covers the five main steps of Q methodology, from defining the concourse to data analysis and interpretation. The program includes research examples from communication and other fields, along with practical exercises in research design. It aims to provide an introductory, critical, and practical understanding of the method. By the end, participants should be able to assess its suitability and apply it in their own academic research.
Content Analysis (6 hours)
12 and 13 February, 2–4 p.m.
Professor Rafael Sampaio (UFPR)
This course introduces students to scientific quantitative content analysis. It presents the main rules for developing codes and analytical categories, including a suitable codebook. It then covers good practices for training coders. Finally, it addresses the importance of intercoder reliability testing and introductory ways to conduct it according to international standards. Within this scope, the course briefly examines the foundations of the technique’s scientific character, including replicability, reliability, and validity.
