Research Question and Background
Housing affordability is a persistent policy concern across Australian capital cities, and Greater Sydney consistently records some of the least affordable housing in the country. Housing stress is conventionally defined using the ratio of housing costs to gross household income, with the widely used benchmark identifying stress where lower-income households spend more than 30 per cent of income on housing. Rather than dichotomise at this threshold, the present analysis models the continuous housing cost-to-income ratio, which preserves information and allows the size of each association to be estimated directly.
The research question is: among households in Greater Sydney, which household and locational characteristics are associated with the proportion of gross income devoted to housing costs? The analysis is intended to inform where affordability pressure concentrates and which household types are most exposed.
Data and Variables
The working dataset contains 1,860 households drawn from a cross-sectional survey spanning the statistical area level 3 regions of Greater Sydney. The unit of analysis is the household. The dependent variable is the housing cost-to-income ratio, expressed in percentage points, calculated as weekly housing costs (rent or mortgage repayments) divided by gross weekly household income and multiplied by 100. The mean ratio was 28.4 percentage points (standard deviation 10.7).
Seven predictors were included:
- Gross weekly household income, in thousands of dollars (continuous).
- Private renter, coded 1 for households renting privately and 0 for owner-occupiers with a mortgage (reference).
- Distance from the central business district, in tens of kilometres (continuous).
- Household size, in usual residents (continuous).
- Apartment, coded 1 for flats or apartments and 0 for separate houses and other dwellings (reference).
- Commonwealth Rent Assistance recipient, coded 1 where the household received the payment and 0 otherwise.
- Single-parent household, coded 1 for one-parent families with dependent children and 0 otherwise.
Before modelling, the distribution of every variable was examined. The mean gross weekly household income was 2.31 thousand dollars (standard deviation 1.12), the median distance from the central business district was 22 kilometres, and the mean household size was 2.7 usual residents. Private renters comprised 41 per cent of the sample, apartment dwellers 29 per cent, Commonwealth Rent Assistance recipients 18 per cent, and single-parent households 11 per cent. Applying the conventional benchmark, 34 per cent of households would be classified as being in housing stress, a proportion broadly consistent with survey estimates for Greater Sydney. No implausible values were detected, and item missingness was below 4 per cent and handled through complete-case analysis.
Analytic Approach
Analyses were performed in Stata version 18. A multiple linear regression (ordinary least squares) was estimated with the housing cost-to-income ratio as the outcome and the seven predictors entered together. Standardised (beta) coefficients were computed to compare the relative strength of predictors measured on different scales. Because a formal test indicated non-constant error variance, the model was re-estimated with heteroscedasticity-robust (Huber and White) standard errors using the vce(robust) option, and the robust results are reported.
Ordinary least squares was chosen because the outcome is continuous and approximately interval-scaled, and because the aim was to estimate and interpret the marginal association of each predictor rather than to maximise prediction. Predictors were entered simultaneously on theoretical grounds rather than selected by a stepwise algorithm, which avoids the biased standard errors and unstable coefficients that automated selection can produce.
Assumption and Diagnostic Checks
The standard assumptions of ordinary least squares were examined. Linearity between each continuous predictor and the outcome was assessed with augmented component-plus-residual plots, which showed no marked departures from linearity. The normality of residuals was inspected using a histogram, a normal quantile plot and the Shapiro and Wilk test; residuals showed mild positive skew, but given the large sample the central limit theorem supports valid inference. Homoscedasticity was tested with the Breusch and Pagan and Cook and Weisberg test, which was significant, chi-square (1) = 41.3, p < .001, indicating heteroscedasticity; robust standard errors were therefore adopted. Multicollinearity was assessed using variance inflation factors, which averaged 1.34 and did not exceed 1.9 for any predictor, so collinearity was not a concern. Influential observations were reviewed using leverage-versus-squared-residual plots, Cook’s distance and DFBETA statistics; a small number of high-leverage cases were examined individually and retained, as none materially altered the coefficients.
Results
The model was statistically significant and explained a substantial share of variation in the housing cost-to-income ratio, F(7, 1852) = 138.2, p < .001, R-squared = .343, adjusted R-squared = .341. The root mean squared error was 8.7 percentage points. Table 1 reports the unstandardised coefficients, robust standard errors, t statistics, p values, confidence intervals and standardised coefficients.
| Predictor | B | Robust SE | t | p | 95% CI | Beta |
|---|---|---|---|---|---|---|
| Gross weekly income ($1,000s) | -5.86 | 0.41 | -14.29 | <.001 | -6.66 to -5.06 | -0.38 |
| Private renter | 4.72 | 0.68 | 6.94 | <.001 | 3.39 to 6.05 | 0.17 |
| Distance from CBD (per 10 km) | -1.58 | 0.27 | -5.85 | <.001 | -2.11 to -1.05 | -0.14 |
| Household size (persons) | 1.34 | 0.31 | 4.32 | <.001 | 0.73 to 1.95 | 0.11 |
| Apartment | -1.92 | 0.62 | -3.10 | .002 | -3.14 to -0.70 | -0.08 |
| Rent Assistance recipient | 3.18 | 0.74 | 4.30 | <.001 | 1.73 to 4.63 | 0.12 |
| Single-parent household | 2.06 | 0.79 | 2.61 | .009 | 0.51 to 3.61 | 0.07 |
| Constant | 41.62 | 1.90 | 21.91 | <.001 | 37.90 to 45.34 |
All seven predictors were statistically significant. Gross household income was by far the strongest determinant: each additional 1,000 dollars of weekly income was associated with a reduction of 5.86 percentage points in the housing cost-to-income ratio, and its standardised coefficient of -0.38 was the largest in absolute terms. Private renters devoted 4.72 more percentage points of income to housing than owner-occupiers with a mortgage, holding other factors constant. Households further from the central business district spent less, with each ten kilometres associated with a 1.58 percentage-point reduction, consistent with the well-documented trade-off between location and price. Larger households, Rent Assistance recipients and single-parent households all showed higher ratios, while apartment dwellers showed slightly lower ratios than those in separate houses.
The standardised coefficients clarify the relative importance of each predictor. Income dominated, with a beta more than twice the magnitude of the next strongest predictor, private renting. Locational and household-structure variables each made smaller but meaningful contributions. Together the seven predictors accounted for 34 per cent of the variance in the housing cost-to-income ratio, which is substantial for cross-sectional household data and leaves the remainder to unmeasured influences such as accumulated wealth, employment stability and individual housing histories.
Interpretation
The results confirm that income is the dominant driver of housing stress, which is intuitive given that the outcome is itself income-normalised, but the analysis also isolates several structural and locational effects net of income. The renter penalty is particularly policy-relevant: even after adjusting for income and location, private renters carry a meaningfully higher housing burden than mortgaged owner-occupiers, which is consistent with concerns about the security and affordability of the private rental market in Sydney. The distance gradient quantifies the affordability trade-off that pushes lower-income households toward the urban fringe, where transport costs and access to services may offset nominal housing savings.
The positive coefficient for Rent Assistance recipients should be read carefully. It does not imply that the payment increases stress; rather, receipt of the payment flags households that are already on low incomes and in the rental market, and the residual stress after the payment suggests the assistance does not fully close the affordability gap for these households. The single-parent and household-size effects point to family structure as an additional axis of vulnerability that operates independently of income.
From a policy standpoint, the analysis suggests that affordability responses cannot rely on income measures alone. The persistence of a renter penalty after adjustment indicates that tenure security and rental supply are distinct levers, while the distance gradient implies that transport and services on the urban fringe form part of the affordability picture rather than a separate concern. Because family structure and household size raise stress independently, targeted assistance for single-parent and larger households may reach vulnerability that broad income thresholds overlook. These implications are descriptive rather than prescriptive, but they help indicate where pressure concentrates.
Limitations
Several limitations apply. The data are cross-sectional, so the associations describe patterns at one point in time and cannot capture how households move into or out of stress. Housing costs were measured as current rent or mortgage repayments and do not include rates, insurance, utilities or the imputed cost of outright ownership, so the ratio understates total housing outlays for some groups. The model omits wealth, savings and access to family support, which may buffer stress independently of income. The presence of heteroscedasticity, although addressed with robust standard errors, indicates that the variance of housing stress differs systematically across households, which a quantile regression could explore more fully. Finally, households were treated as independent, whereas clustering within statistical areas may induce correlated errors; a multilevel model with region-level random effects would be a natural extension. A further extension would model the endogeneity of housing tenure, since the decision to rent or to buy is itself shaped by income and life stage, which a single-equation specification cannot fully disentangle.
References
Australian Bureau of Statistics. (2022). Housing occupancy and costs, Australia, 2019 to 2020. Australian Bureau of Statistics.
Australian Housing and Urban Research Institute. (2021). Understanding rental stress in Australian cities. AHURI Final Report Series.
Gabriel, M., Jacobs, K., Arthurson, K., & Burke, T. (2019). Conceptualising and measuring the housing affordability problem. Housing Studies, 34(6), 991 to 1010.
Hulse, K., Reynolds, M., & Yates, J. (2020). Changes in the supply of affordable housing in Australian cities. Urban Studies, 57(9), 1820 to 1838.
Long, J. S., & Freese, J. (2014). Regression models for categorical dependent variables using Stata (3rd ed.). Stata Press.
Productivity Commission. (2019). Vulnerable private renters: Evidence and options. Productivity Commission.
Rowley, S., & Ong, R. (2018). Housing affordability, stress and low-income households in Australia. Australian Economic Review, 51(3), 358 to 372.
Wooldridge, J. M. (2019). Introductory econometrics: A modern approach (7th ed.). Cengage Learning.
Yates, J. (2016). Why does Australia have an affordable housing problem and what can be done about it? Australian Economic Review, 49(3), 328 to 339.