Does Gender Predict Judicial Agreement on Canada's Supreme Court?
Measuring Correlations between Agreement and Identity Markers
This is an AI generated essay that explains the code used to make this YouTube video.
Judges arrive at the Supreme Court of Canada with different personal histories and professional experiences. Some spent most of their careers representing governments or large institutions, while others represented individuals. Some came from civil law, others from common law. They attended different law schools, began judging at different ages, practised in different fields, and were appointed by different prime ministers.
Do any of these differences help explain which judges agree with one another?
This study approached that question by treating the pair of judges, rather than the individual judge, as the unit of analysis. Instead of asking whether women vote differently from men, for example, it asked whether two judges of the same sex agree more often than a pair consisting of a man and a woman. The same logic was applied to province of origin, law school, primary language, appointing party, legal background, length of practice and numerous other characteristics.
The objective was not to identify what “causes” a judge to decide a case in a particular way. Judicial decisions emerge from the facts, legal questions, precedents, arguments and institutional setting of each case. The narrower objective was to determine whether judges who resemble one another in certain ways also tend to vote together more often—and whether any observed pattern was stronger than could reasonably be expected by chance.
Constructing the judge pairs
The analysis included 68 pairs of Supreme Court of Canada judges and tested 25 predictors. For every pair, the study calculated three forms of agreement:
1. Opinion-group agreement: whether the two judges joined the same set of reasons.
2. Disposition-side agreement: whether they were on the same side of the Court’s disposition.
3. Outcome agreement: whether they supported the same ultimate result.
These measures capture related but distinguishable forms of agreement. Two judges may support the same result while writing or joining different reasons. Opinion-group agreement is therefore the most demanding measure, while outcome agreement captures the broader question of whether the judges ultimately voted the same way.
Pairs did not necessarily hear the same number of cases together. A pair that sat together in many cases provides a more stable estimate of agreement than one that shared only a few cases. The analysis therefore weighted each pair according to the number of cases for which the relevant agreement measure was known. A pair with more shared cases had more influence on the estimate than a pair with fewer shared cases.
Turning judicial characteristics into pairwise predictors
The predictors fell into two main categories.
For categorical characteristics, the analysis asked whether the two judges shared the same value. Examples included:
· Were they of the same sex?
· Did they attend the same law school?
· Were they appointed from the same province?
· Did they have the same primary language?
· Were they appointed by the same prime minister?
· Were they trained in the same legal tradition?
Each pair received a value of one when the characteristic matched and zero when it did not. The principal effect was then calculated as:
agreement among matching pairs minus agreement among non-matching pairs.
Suppose same-sex pairs had an average agreement rate of 75 per cent and mixed-sex pairs had an average rate of 71 per cent. The estimated effect of sharing the same sex would be four percentage points.
Continuous characteristics required a different approach. For variables such as age, year of birth or years spent practising before becoming a judge, the study calculated the absolute difference between the two judges. It then estimated a weighted linear regression relating that difference to their agreement rate.
For these variables, a negative slope means that judges who were more similar tended to agree more. A slope of −0.011, for example, means that every additional unit of difference—one year, in this case—was associated with an agreement rate approximately 1.1 percentage points lower.
The analysis also reported a weighted R² for each predictor. R² describes how much of the variation in pairwise agreement is statistically associated with that one variable. It should not be read as proof that the predictor caused the variation, particularly because the models tested each predictor separately rather than estimating a single causal model containing all of them.
Testing the patterns through permutation
An observed association is not automatically meaningful. With many predictors and a relatively small collection of judge pairs, some patterns will appear simply by accident.
To assess this possibility, the study used a permutation test. For each predictor, the judges’ metadata values were randomly reassigned among the judges. The pairwise characteristic was then reconstructed and its relationship with agreement recalculated.
This procedure preserved several important features of the data:
· the real agreement rates;
· the actual network of judge pairs;
· the number of cases attached to each pair; and
· the overall distribution of the predictor.
What it destroyed was the link between a particular judge and that judge’s actual characteristic. For example, the same number of men and women remained in the dataset, but the sex labels were randomly reassigned to different judges.
This process was repeated 5,000 times for every predictor and each agreement measure. The actual result could then be compared with the distribution of results produced under random reassignment. The directional permutation p-value represents the proportion of randomized datasets that produced an effect at least as strong as the one observed in the expected direction.
For a shared categorical characteristic, the expected direction was positive: judges sharing that characteristic were hypothesized to agree more. For continuous variables, the expected direction was negative: judges closer together were hypothesized to agree more.
Negative controls
The study also included deliberately arbitrary predictors that should have no plausible connection to judicial reasoning. These included:
· whether a surname began with A–M or N–Z;
· whether a judge was appointed to the Supreme Court in an even or odd year;
· surname length;
· alphabetical surname rank;
· last-name initial; and
· birth year divided into arbitrary remainder categories.
These negative controls provide a useful reality check. If a genuine biographical predictor produces a pattern no stronger than patterns based on surname length or appointment-year parity, it becomes harder to argue that the apparent relationship reflects something substantive.
Negative controls are especially important here because the dataset contains only 68 pairs, the same judges appear in several different pairs, and many predictors were tested. Under those conditions, conventional statistical thresholds can create a false sense of certainty.
The clearest result: similarity in pre-judicial practice
The most consistent substantive finding concerned the number of years each judge spent practising law before receiving a first judicial appointment.
Across all three definitions of agreement, judges with similar amounts of pre-judicial practice tended to agree more:
· For opinion-group agreement, the slope was −0.0110, with a directional permutation p-value of 0.0154and a weighted R² of 0.279.
· For disposition-side agreement, the slope was −0.0099, with a p-value of 0.0174 and an R² of 0.315.
· For outcome agreement, the slope was −0.0098, with a p-value of 0.0210 and an R² of 0.363.
In practical terms, a ten-year difference in the amount of time two judges practised before joining the bench was associated with an agreement rate roughly 10 percentage points lower, according to the fitted one-predictor models.
The consistency of the finding is notable. It appeared under the narrowest definition of agreement, under the broadest definition and under the intermediate measure. Its permutation p-values also fell below 0.05 in all three analyses.
One possible interpretation is that lawyers who spend similar lengths of time in practice develop comparable professional instincts. The amount of time spent as counsel may shape how lawyers understand evidence, litigation strategy, institutional constraints and the practical consequences of legal rules. Lawyers appointed early may carry a different mixture of professional habits onto the bench than those who spent decades in practice.
That explanation remains speculative. Years of practice may also stand in for other factors, such as generation, career trajectory, professional specialization or the historical period in which a judge entered the judiciary. The study identifies a relationship; it does not establish its mechanism.
Sex, law school and other suggestive patterns
Several other characteristics produced suggestive but weaker associations.
Same-sex pairs had opinion-group agreement approximately 4.0 percentage points higher than mixed-sex pairs. The corresponding directional permutation p-value was 0.0836, and the predictor accounted for about 3 per cent of the weighted variation in pairwise opinion agreement.
For disposition-side agreement, the same-sex difference was approximately 3.1 percentage points, with a p-value of 0.0816. For outcome agreement, it fell to approximately 2.1 percentage points, with a p-value of 0.1140.
These results do not justify the simple conclusion that male and female judges vote differently. The comparison is between same-sex and mixed-sex pairs, not between the average votes of male and female judges. More importantly, the observed differences did not meet the conventional 0.05 threshold in the permutation tests and explained relatively little variation.
Law school produced a somewhat larger estimated difference. Judges who attended the same law school had opinion-group agreement approximately 7.1 percentage points higher than judges from different schools. The corresponding p-value was 0.1036. For outcome agreement, the difference was approximately 6.4 percentage points, with a p-value of 0.0526—close to, but still above, the conventional threshold.
This could reflect shared legal training, professional networks or regional background. It could also be an unstable result driven by the small number of same-school pairs or by particular judges who both attended the same institution. The analysis does not allow those possibilities to be cleanly separated.
Similarity in years spent on the bench before appointment to the Supreme Court also pointed in the expected direction. Judges with more similar amounts of prior judicial experience tended to agree somewhat more, but its permutation p-values ranged from approximately 0.109 to 0.117 and its R² values were around 4 to 5 per cent.
Primary language and legal tradition showed small positive effects. Province of origin, appointing prime minister and appointing party showed little relationship with agreement. In particular, whether two judges were appointed by governments of the same political party produced effects close to zero under all three agreement measures.
Findings that ran against the hypothesis
Not all predictors pointed in the expected direction.
Pairs sharing the same race or ethnic classification had lower estimated agreement than pairs assigned to different categories. Pairs sharing the same pre-judicial area of law also had lower agreement—approximately six to seven percentage points lower for some outcomes.
These results had large directional p-values because the original hypothesis predicted that similarity would increase agreement. They should not be interpreted as strong evidence that judges from different racial backgrounds or different areas of practice are inherently more likely to agree. Categorical variables of this kind compress complicated biographies into coarse labels, and some categories may contain very few judges or pairs.
A “same area of law” variable, for example, treats every mismatch alike. A criminal lawyer paired with a commercial lawyer is coded the same way as a constitutional lawyer paired with an administrative lawyer. It also ignores judges with broad or changing practices. The negative result may therefore reflect limitations in classification rather than a meaningful reverse relationship.
What the negative controls revealed
The negative controls counsel against overconfidence.
The arbitrary variable based on whether a judge was appointed in an even or odd year produced a same-category advantage of approximately five percentage points. It reached a directional permutation p-value below 0.05 for disposition-side and outcome agreement and came close for opinion-group agreement.
There is no persuasive theory under which even-numbered appointment years should make two judges reason alike. Its apparent success demonstrates that a small dataset can generate convincing-looking patterns from meaningless variables.
The study’s formal negative-control comparison therefore reached an important conclusion: no real predictor simultaneously achieved a directional p-value of 0.05 or less and exceeded every negative control of the same type by absolute effect size.
This does not erase the result for years of pre-judicial practice. That predictor was more consistent across outcomes, had substantially larger R² values and is supported by a plausible professional mechanism. It does mean that the finding should be treated as a strong exploratory signal rather than a settled discovery.
Limitations
Several limitations shape what can be concluded.
First, the study contains only 68 judge pairs. Moreover, these are not 68 fully independent observations. The same judge appears in many pairs, meaning that a particularly distinctive judge can affect several observations at once.
Second, the predictors were evaluated one at a time. Many judicial characteristics overlap. Judges of similar ages may have been called to the bar at similar times, appointed by the same governments, spent similar numbers of years in practice and served together during the same period. A large one-predictor R² does not show that the variable has an independent effect after accounting for these relationships.
Third, 25 predictors were tested against three outcomes. That creates many opportunities for chance findings. The permutation tests evaluate each predictor against its own randomized distribution, but the reported p-values were not adjusted for the full number of comparisons.
Fourth, some categorical characteristics are difficult to define and may contain small or uneven groups. “Area of law,” “race or ethnic origin,” “represented institutions or individuals,” and even “primary language” can flatten complicated careers and identities into a single label.
Fifth, pairwise agreement is partly determined by case assignment and Court composition. Two judges may appear especially similar because they sat together on a particular mixture of cases. Weighting by the number of shared cases improves the stability of the agreement estimates, but it does not completely remove case-selection effects.
Finally, the study is observational. It can reveal associations between biographical similarity and voting agreement, but it cannot establish that one caused the other.
Conclusion
The broad result of this study is not that judges vote according to sex, political appointment or regional identity. Most of those characteristics had small, inconsistent or statistically uncertain relationships with pairwise agreement.
The strongest pattern was professional rather than conventionally demographic or political. Judges who had spent similar lengths of time practising law before first joining the bench tended to agree more often once they reached the Supreme Court. The relationship appeared under all three definitions of agreement and was unusually large compared with the other predictors.
At the same time, an arbitrary negative control based on even and odd appointment years also generated a superficially significant pattern. That result is a warning against treating any single p-value as decisive.
The most defensible conclusion is therefore provisional: the length and timing of a judge’s professional formation may help explain judicial alignment, but the available dataset is too small and interconnected to establish this confidently. The study identifies a promising hypothesis for further investigation rather than a final theory of judicial behaviour.
A larger study could test the same question across a longer period of Supreme Court history, incorporate case-level controls, account explicitly for repeated appearances by the same judges and estimate several predictors together. Until then, the results are best understood as a map of where stronger evidence may be found—not as proof that a judge’s biography determines how that judge will vote.

