The 16th Spanish Stata Conference will take place on 15 October 2026.
The conference will feature invited sessions from current StataCorp developers as well as an optional drinks reception and dinner. Don't miss this opportunity to learn new and exciting applications of Stata, engage with StataCorp's developers, and network with researchers from across all disciplines. Because of the international nature of the conference, all presentations will be held in English.
All times are in CEST (UTC +2)
| 8:30–9:00 | Welcome and registration |
| 9:00–10:00 | dime and netreg: Interactive network graphs of marginal effects
Abstract:
Graphs are widely used to represent social structures and to study relationships between variables. This presentation proposes an approach that extends their analytical potential by combining regression modeling with interactive network visualization. The method fits generalized linear models (including multinomial models) and their mixed-effects extensions.
Modesto Escobar
|
| 10:00–10:30 | ages: A new community-contributed command to estimate the impact of the population's age structure
Abstract:
ages is a new community-contributed module that estimates the relationship between different population-age groups (cohorts) and a dependent variable of choice, based on the polynomial transformation suggested by Fair and Dominguez (1991) and applied later by Higgins (1998), Arnott and Chaves (2012), and Juselius and Takas (2021). Capturing the effect of the age structure in a regression involves different methodological issues, such as weak precision if the number of population cohorts is large compared with the number of time periods, large correlation between consecutive cohorts’ shares and perfect collinearity with respect to the constant. The solution consists of restricting the age groups' coefficients to lie on a Pth degree polynomial and restricting its sum to be equal to zero to remove the perfect multicolinearity. The ages command performs all the necessary transformations and mathematical calculations to estimate and report the final effect by age cohort, together with its standard error and confidence interval, as an output matrix and through an automated graph that displays the impact of age cohort on the dependent variable. It allows testing for different age structure effects based on different groups captured by an interaction with a dummy variable. It also allows estimating a fully dynamic model, allowing the user to define the desired number of lags of the age structure groups, reporting the age structure by each lag and as sum of all them. ages also allows the user to choose the preferred time-series or panel-data method. The user can easily choose different options for the graph with the impact by age group, as well as several other options to save and use the results.
Alfonso Ugarte-Ruiz
|
| 10:30–10:45 | Break |
| 10:45–11:15 | Target trial emulation in Stata
Abstract:
This presentation will demonstrate, through an example in Stata, how to conduct a target trial emulation. Target trial emulation is a practical framework that uses observational data to answer causal questions by mimicking a hypothetical randomized clinical trial. In this context, it is essential to control for potential biases and confounding through causal inference methods in order to correctly estimate the causal effect of treatments on the event of interest. Some of the methods that will be discussed and implemented in Stata during the presentation include inverse probability of treatment weighting (IPTW) and the g-formula. Using the comparison of two treatments as an example, the necessary steps to implement an emulated clinical trial will be presented. First, the data will be structured and expanded into a weekly format. Subsequently, different methods will be applied to estimate the causal effect, including a pooled logistic regression model, IPTW, the estimation of adjusted survival curves, and the g-formula. The results will be evaluated under different scenarios to assess how changes in the analytical assumptions affect the findings. The ultimate aim is for attendees to understand how to translate a clinical question into a reproducible, transparent Stata analysis that is consistent with the principles of causal inference.
Elisa Medrano Buendíam
|
| 11:15–11:45 | netcorr: Interactive graphs for correlations and coincidences
Abstract:
Exploratory graphical analysis remains essential in an era of abundant data and increasingly sophisticated data science methods. Since Tukey's seminal work.
Cristina Calvo Lopez
|
| 11:45–12:15 | The socioeconomic determinants of fertility
Abstract:
Mozambique has recorded a decline in its overall fertility rate, an important element of the demographic transition and the exploitation of the demographic dividend. In fact, while in 1997, the country average was estimated at 5.2 children per woman, in 2017 available estimates indicated 4.9 children per woman. The reduction in the fertility rate has, however, not been homogeneous across the country. While in urban areas women have on average 3.6 children, in rural areas this rate remains high at 5.8 children per woman. By region and province, the disparity in the fertility rate is even more pronounced: the southern region registers the lowest rates, with the extreme case of Maputo City, where each woman has an average of only 2.8 children. Conversely, the northern region has the highest fertility rates, with the extreme case of the provinces of Niassa and Cabo Delgado, where the fertility rate is 6.8 and 6.2 children per woman, respectively. The fertility rate in the province of Nampula is even higher than the rest of the country's provinces, at 5.8 children per woman. It can therefore be stated that one of the greatest challenges to achieving a sustainable demographic dividend in the country lies in influencing adequate changes in demographic behavior outside the south of the country, with an even more intense focus on the northern region. The communities in the north and center of the country are essentially traditional. This study aims to understand the socioeconomic factors that influence the difference in fertility between the three regions of the country in such a way that the most important elements on which to focus for a more accelerated reduction can be identified.
Maimuna Assiate Ibraimo
|
| 12:15–1:15 | Lunch |
| 1:15–2:15 | Introduction to panel VAR models using Stata
Abstract:
Panel vector autoregression (panel VAR) generalizes dynamic panel models and time-series VAR estimation by modeling the time dynamics while accounting for unobserved cross-sectional heterogeneity. This presentation provides an end-to-end empirical workflow for implementing a panel VAR model with real-world data. I will discuss model formulation, lag selection, and managing instrument proliferation from excessive moment conditions in GMM estimation. I will then demonstrate how to assess model stability and interpret impulse–response functions (IRFs) to evaluate the impact of structural shocks on the endogenous variables in the system.
Gustavo Sánchez
StataCorp
|
| 2:15–2:35 | Beyond the AUC: A comprehensive workflow in Stata for evaluating clinical prediction models—application to a multicenter cardiovascular risk model
Abstract:
With the increasing application of clinical prediction models in medical research, evaluating model performance has become an essential component of the development and validation process. However, most studies still rely primarily on the area under the ROC curve (AUC) to assess model performance, paying less attention to key indicators such as calibration, incremental model improvement, clinical value, and external generalizability. Guidelines for evaluating predictive models, such as TRIPOD, recommend a multidimensional assessment, yet in Stata, these analyses typically require multiple independent commands.
Shengan Li
|
| 2:35–2:55 | Does your covariance structure choice matter? Evidence from a systematic Monte Carlo study with mixed
Abstract:
Stata’s mixed command offers three common random-effects covariance structures (independent, exchangeable, unstructured) with little guidance on which to choose or what a wrong choice actually costs. This study analyzes results from a systematic Monte Carlo study varying the true correlation among random effects, the heterogeneity of their variances, cluster count and size, and the number of random effects to answer that question directly. Four findings carry direct implications for practice. Fixed-effect point estimates are robust to structure choice, but exchangeable can miscalibrate confidence-interval coverage even when its coefficients look unaffected, so a check comparing only point estimates can miss the problem. Independent biases the covariances by construction but leaves variances close to unbiased even under strong true correlation, so its cost is often smaller than assumed. Exchangeable can be actively misleading under heterogeneity, in one scenario reversing a covariance's sign entirely. Unstructured convergence degrades sharply as random effects grow relative to cluster count. The study derives concrete decision rules for when the extra unstructured parameters are worth fitting.
Antonio M Jaime Castillo
|
| 2:55–3:15 | surveye: Create an interactive HTML dashboard from a Survey Solutions or SurveyCTO questionnaire and the corresponding Stata data
Abstract:
surveye converts a Survey Solutions or SurveyCTO questionnaire HTML file and the corresponding Stata dataset into a polished, standalone, interactive HTML dashboard. It automatically uses the questionnaire structure, wording, response categories, and section order to organize the output. The command supports filters, highlights, custom variables, related-variable groups, subgroup comparisons, weighted and unweighted estimates, local-currency/USD switching, profile tables, numeric summaries, optional confidence intervals, customizable themes, right-to-left Arabic and Urdu interfaces, and optional GPS maps. The dataset in memory is preserved, and all required JavaScript, styles, and selected data are embedded in the generated dashboard. The package includes its Java engine and requires no separate Stata dependencies.
Attique Ur Rehman
|
| 3:25–3:55 | Coca as a subsistence good: Displacement of food expenditure and habitual consumption in Bolivia and Peru (2006–2024)
Abstract:
Coca leaf chewing (acullico) is typically viewed as a cultural custom. This presentation—which extends the evidence documented for Bolivia by Gutierrez Miranda (2023) to Peru and a two-country comparative framework—argues, based on two decades of microdata, that it is also an economic phenomenon: coca functions as a subsistence good for poor households, and its satiating effect—physiologically documented yet lacking real nutritional value (Bolton, 1976; Penny et al., 2009)—crowds out spending on actual, more expensive food. The analysis is framed within the economics of addictive consumption and temptation goods (Becker and Murphy 1988; Banerjee and Mullainathan 2010). A harmonized two-country panel was constructed (Bolivia household survey and Peru's ENAHO, 2006–2024; ˜1.13 million household-years), made comparable regarding poverty and income through relative position and PPP-adjusted lines. The empirical strategy combines a pooled probit model (participation), a two-part/double-hurdle model (intensity), and the actual reported ordinal consumption frequency—derived from expenditure modules to account for self-supply bias. Results show that coca prevalence falls monotonically with income (Bolivia: 48% → 23%; Peru: 17% → 2%), whereas alcohol consumption does not; the indigenous effect remains robust when controlling for income and education. The subsistence gradient operates at the extensive margin (who consumes) rather than the frequency margin; consumption frequency is habitual and consistent across consumers—unlike alcohol consumption, which is social and sporadic. The methodological contribution consists of a reproducible R–Stata workflow (estimation in R; validation, margins, twopm, and graphing in Stata, ensuring exact replication) and a community-contributed command (cruceandino) that generates both the table (table and collect) and the figure for a multiway cross-tabulation in a single step—designed for consumption studies using household surveys and proposed for the SSC Archive repository.
Juan Marcelo Gutierrez Miranda
|
| 3:55–4:10 | Break |
| 4:10–4:25 | Synthetic indicators of the 2030 Agenda: Evidence from Eurostat using Stata
Abstract:
This presentation constructs synthetic sustainability indicators based on the 2030 Agenda using harmonized Eurostat data and fully reproducible procedures implemented entirely in Stata. I consider time series for the 169 targets of the sustainable development goals (SDGs) over the period 2000–2023, covering 27 European countries: Austria, Belgium, Bulgaria, Croatia, Cyprus, Czech Republic, Denmark, Estonia, Finland, France, Greece, Germany, Hungary, Ireland, Italy, Latvia, Lithuania, Luxembourg, Malta, the Netherlands, Poland, Portugal, Romania, Slovenia, Slovak Republic, Spain, and Sweden. The indicators are downloaded directly from Eurostat using the community-contributed eurostatuse command, ensuring transparency, traceability, and easy data updates. Based on these series, synthetic indices are constructed following the methodology proposed by Rivero and Fernández (2008). The SDG targets are normalized on a scale from 0 to 1, distinguishing between direct and inverse variables, so that higher values reflect a greater probability of achieving the corresponding goal. Subsequently, the normalized values are aggregated to obtain partial indices for each SDG, calculated as the average of the targets included in each goal. Following the classification of the SDGs into five pillars proposed by Tremblay et al. (2020), the goal-level indices are weighted and aggregated to construct synthetic indicators for each pillar, in line with the institutional structure of the 2030 Agenda. The entire process of normalization, weighting, and aggregation is implemented in Stata, allowing for a coherent treatment of panel data and facilitating the replicability of the analysis.
Building on these synthetic indicators, the presentation conducts two complementary empirical applications. First, this systematic treatment allow us to examine the causal interactions among the 5 pillars of the 2030 Agenda for 28 countries over the period 2000–2021. The results reveal a leadership role of the planet pillar over the remaining pillars, while the planet, peace, prosperity, and partnership pillars jointly exhibit leadership over the people pillar. These findings highlight the central importance of environmental sustainability and international cooperation as prerequisites for social well-being and economic prosperity. Second, this novel dataset also assesses the efficiency of public spending in achieving the SDGs in 27 European countries during the period 1995–2023 by applying data envelopment analysis (DEA) models. The results show efficiency scores ranging from 0.70 to 0.94 for inputs and from 0.80 to 0.94 for outputs, indicating potential improvement margins of up to 30.6% in resource use or 20.3% in outcomes. Denmark, Ireland, Finland, and Sweden emerge as consistent benchmarks of sustained efficiency over time. Overall, the combined use of eurostatuse and customized Stata routines provides a flexible and fully reproducible framework for the comparative analysis of sustainability performance and policy efficiency in Europe.
Najat Bazah Lamchanna
|
| 4:25–4:40 | Mechanisms linking individual education and european identification: A case study of Spanish youth
Abstract:
A robust literature documents that education is one of the strongest individual-level predictors of European identification (EUID), yet the mechanisms underlying this relationship remain incompletely understood. Standard cross-national surveys capture EUID but lack the detailed measures needed to test key theorized pathways, particularly the transactionalist expectation that transnational travel experiences shape identification. Drawing on YOEDER (youth education and European attitudes), a 4-wave panel survey of Spanish youth aged 18–29 (N = 16,073), this presentation estimates a multiple-indicators multiple-causes (MIMIC) structural equation model to decompose the education–EUID link through 7 mediators: EU travel and satisfaction, EU institutional knowledge, universalist values, income, press consumption, peers' opinions of the EU, and foreign language ability. Results show that education is positively associated with EUID and six of the seven mediators and that all mediators are positively linked to EUID. Satisfactory travel experience within the EU emerges as the dominant channel, accounting for roughly 44% of education's total effect, and this pattern holds under an alternative sequential specification and after accounting for travel to other world regions. Disentangling frequency from satisfaction shows that the affective quality of travel, not mere exposure, drives the mediation. These findings extend transactionalist accounts of European identity by highlighting the concrete experiential pathway through which education fosters identification with Europe.
Juan J Fernández
|
| 4:45–5:05 | Household shocks, resilience, and children’s educational outcomes: Evidence from rural Punjab, Pakistan
Abstract:
Household shocks are a pervasive form of disruption to children’s education in developing countries, yet the extent to which their consequences depend on the nature of the shock remains poorly understood. This presentation studies two of the most common shocks to rural households—parental illness and harvest loss—both plausibly income-reducing but theoretically expected to operate through different mechanisms. Using four rounds of the learning and educational achievement in Punjab schools (LEAPS) data from rural Punjab, Pakistan, I find that these different channels appear to drive the observed differential effects of parental illness and harvest loss on children’s standardized test scores. Father’s illness is consistently associated with lower test scores: a one-standard-deviation increase in the father’s health impairment predicts a decline in mathematics and Urdu performance, equivalent to approximately one month of learning progress documented in the LEAPS sample (Bau et al. 2021). The effects are concentrated at moderate illness severity and are greater for boys than girls. I find no such corresponding effects for the mother’s illness.
In contrast, negative harvest shocks are not associated with a detectable effect on test scores across two-way fixed-effects, instrumental-variable, and staggered difference-in-differences specifications. I find no suggestive evidence that this average null result conceals heterogeneous subgroup effects: an exploratory causal forest vulnerability analysis finds no subpopulation of children experiencing larger learning losses from harvest shocks. For father’s illness, however, the evidence points to a vulnerability pattern—children in low-asset households with thin social networks experience estimated effects nearly 1.5 times the sample average, while low-vulnerability children exhibit no significant effect. These findings suggest that the nature of the shock matters: parental illness appears more strongly linked to children’s education than harvest shocks, and its educational costs are disproportionately concentrated among already disadvantaged households.
Jovera Shakeel
|
| 5:05–5:35 | Open panel discussion with Stata developers
Contribute to the Stata community by sharing your feedback with StataCorp's developers. From feature improvements to bug fixes and new ways to analyze data, we want to hear how Stata can be made better for our users.
|
Gustavo Sánchez
Gustavo Sánchez is a Senior Econometrician and Director of the Technical Services department at StataCorp LLC. He has a master's degree in econometrics from Southampton University, UK, and he got his PhD in agricultural economics at Texas A&M University. Gustavo worked at the Central Bank of Venezuela, and he was a professor of econometrics at the Universidad Central de Venezuela.
Isabel Canette
Isabel Canette is the Associate Director of Technical Services and Principal Mathematician and Statistician at StataCorp LLC. She has a PhD in mathematics from Universidad de la República, Uruguay. She has worked for StataCorp for more than 15 years, and she enjoys interacting with Stata users from all over the world and with diverse backgrounds. She performs research for different developement projects, mainly in the areas of survival analysis and multilevel models.
Registration information forthcoming.
Visit the official conference page for more information.
The logistics organizer for the 2026 Spanish Stata Conference is Timberlake Consulting S.L., the Stata distributor for Spain.
View the proceedings of previous Stata Conferences and international meetings.