Tuesday, 28 April 2015

HadISD v1.0.3.2014f released

The latest version of HadISD has been made available on the hadobs website.  This version (v1.0.3.2014f) supercedes the preliminary version from earlier this year (v1.0.3.2014p).  There have been further updates to the ISD source data for the years 2012,2013 and 2014 since the preliminary dataset was created in January, but no changes in earlier years.  

The raw data were downloaded on 7th April 2014, and processed over the subsequent days.  Despite the updates to the ISD, as the previous version was preliminary, we retain the version number (v1.0.3.2014) and only increment the descriptor to "final". This version still contains 6103 stations, with 4060 passing the final filtering checks, as for the preliminary version.  

As always, if you find anything untoward in the data, please contact the dataset maintainers.

Monday, 19 January 2015

v1.0.3.2014p Released

HadISD version 1.0.3.2014p has just been released.  All plots and files should be on the website This update extends the coverage of the dataset to the end of 2014 (31 December at 2300 inclusive).  It remains a preliminary dataset as there could still be further updates to the ISD dataset in the next few months.  We hope to do a processing run for the final version some time around Easter (to create 1.0.3.2014f). 

The raw data were downloaded on 5th January 2014, and processed over the subsequent days.  There have been changes to all of the raw files only in 2013 as part of the normal ISD update process  We have made no substantial changes to the codes which do the conversion to NetCDF files or the Quality Control suite.  Hence the version number has only incremented by 0.0.1 and the year. 

This version still contains 6103 stations, with 4060 passing the final filtering checks, down slightly from the 4071 in v1.0.2.2013p (see the HadISD paper Section 6).  The patterns of flagging are very similar to v1.0.2.2013p (see figures here).  However if you find something strange, do let us know using the contact details on the HadISD website.  Please note the stations which are known to have issues are documented on this blog and on the website.

Fig.1 The fraction of temperature records flagged for each station.

Fig. 2 The fraction of all dewpoint temperature records flagged for each station

Fig. 3 The fraction of all sea-level pressure records flagged for each station
The Homogeneity information for this version is also available on the website using the same procedure (PHA) as outlined in Dunn et al, 2014.

As always, if you see anything untoward in the data or are having problems using it, please do not hesitate to get in touch.

Wednesday, 7 January 2015

Attempting to fix undocumented merges

As mentioned in an earlier posts, we had found some issues in the Canadian stations which appeared like undocumented station moves.  In discussions with Environment Canada, we were given a list of the Canadian WMO stations along with dates of their changes.  There were 994 stations present in their list.  

We separated the stations out into different categories (the number of stations in each is given in parentheses): 

  • Single - stations which appeared in the list only once (529)
  • On/Off - stations which had an "active" and "inactive" status indicating the start and end dates of operation (47)
  • Good Station Moves - stations which showed a change in location, with dates showing the end of reporting at the previous location, and the start in the new location (216)
  • Overlap Moves - similarly to good station moves, but the start of reporting in the new location occurs before the end of reporting at the old (15)
  • Possible Homogeneity issues - multiple dates at a single location indicating perhaps changes in instrumentation (92)
  • Questionable Moves - location changes with no dates given showing the end at one or the beginning at another location (33)
  • Dates - cases where "active" and "inactive" statuses occurred at the same time, so the final status could not be determined (49)
  • Other - more complex sets of start and end dates that could not be categorised easily (13)
In the ISD, there are more than 1000 stations listed as being in Canada.  We selected those which were likely to correspond to the WMO stations (those which have IDs that match 71???0-99999).  This resulted in 934 stations which we could compare to the Environment Canada list.

Stations which appeared in the Single, On/Off and Homogeneity issues categories were retained in the candidate station list.  Those from the Questionable Moves, Dates, Overlap moves and Other were rejected from the station list. 

The 216 stations in the Good Moves list were processed further.  Using the station details in the ISD list, the period of time when the station was in this location as determined from the Environment Canada list was extracted.  Usually this was the most recent location.  The start and end times of the station were adjusted as appropriate to ensure that only the period in the location as given in the full ISD station list was used when further selecting stations.  In many cases this will result in the station not being selected for inclusion with HadISD.

Of the 934 Canadian stations we were able to assess, 797 were kept for processing by further selection criteria, 33 could not be tested and 104 were rejected. 

There are other stations which are located in Canada (which do not match the WMO IDs) which we could not process.  These, along with the 33 which were not in the Environment Canada list, were retained in the stations selection procedure as we have no information indicating that there are problems with them.

These changes result in 14762 stations being selected using the restrictions on latitude, longitude and time-spans, 8561 in the master-list and 8104 in the final merged list (of which 2045 have other stations merged into them).  This a reduction from the 8207 stations which were in the previous selection, but hopefully fewer of these have serious inhomogeneities resulting from the undocumented station moves.

Monday, 20 October 2014

Further thoughts on the Merging Problem

To follow on from the previous post, we've done some more thinking about the issue of merging stations correctly.  The three options from last time were:
  1. We can not merge at all, keep all the ISD station IDs as unique and be confident that by creating HadISD we have not degraded any of the data.
  2. We can merge only in cases where we have specific information (from a national met service, for example) as to the identity of stations.  This could also be applied in cases where we have information indicating that a split of a station record would be appropriate 
  3. We can merge (and split) when we have specific information, but also run an automated procedure to identify candidate stations to merge together.  An example algorithm has been produced by the International Surface Temperature Initiative databank v1.0 (see description in the paper).  However this approach is very likely to introduce some spurious mergers, however careful we are with the algorithm.
with our preference at the time leaning towards number 2.  

When we were thinking of option 3, what we had in mind was something similar to the merging process carried out for HadISD v1.0.0.  In that process the merging of short records was carried out before stations were selected, so that lots of short records, once merged, would be long enough to pass the selection criteria.  This cross-matching of all 29,000 stations in ISD to obtain a parent-set from which HadISD is drawn will result in some final stations being composed of many short segments.  The likelihood that some of these stations will be erroneously merged together is quite high given the automated nature of the build.  A subtly different alternative came to mind.

However, what could be done is to select stations on the raw ISD record lengths and reporting intervals, and then using this master-list, see which other stations in the ISD could be merged in to supplement these primary stations.  This will not increase the final station list (in fact it will decrease it as there will be stations that have been selected that will be merged together), but should improve the data coverage over time for the final set of merged stations.

Fig. 1: Flowchart showing envisaged station selection procedure with merging

Most of the merging process takes place after stations have been selected on the basis of their length of record and their reporting interval, however for specific countries, it occurs before as we have extra and definitive information as to which stations should be merged or split.  At the moment there are also stations in the master list which will be selected to merge with other stations in the master list - hence the reduction to 8207.

Selecting stations to Merge

To select stations which are possible merging candidates we so far are using a very simple algorithm.  We test the horizontal and vertical separation of the stations and also the similarity of the station names.  The distances are mapped to an exponential decay curve, which returns a value between 0 and 1 which we use as a probability.  For the horizontal distances, this curve falls to 1/e by 25km, and by 100m for the vertical separation.  To calculate the similarity of the station names, we use the Jaccard Index (also used by ISTI), also returning a value between 0 and 1.  These three probabilities are multiplied together, and stations where the final value is >0.5 are selected.

For an automated system there is no perfect result - there will always be false positives (stations that we shouldn't merge and are distinct) and false negatives (stations that we should merge but do not select to do so).  An inspection of the resulting candidates for the UK (where we have more idea of the suitability of merging these candidates) suggests that the values used above are reasonable, with no obvious false positives.  The algorithm and thresholds are not yet set in stone, so changes can still occur.

At the current moment in time we find that within the primary station list, 478 stations are similar to others also in the list, reducing the station number to 8207 (including the changes from the German stations outlined below).  By cross matching these 8207 stations to the complete ISD database, 2101 will contain data from other station IDs.

Fig. 2: The effect of the current merging system on the number of stations that report over time.  Improvements at the beginning and end of the record are clearly visible.
Fig. 2 shows the effect of the merging selection as it currently stands on the stations available in each year.  There are clear improvements in the number of merged stations available in the early part of the record (1935-1970) and also the last 10 years or so. 
 

Specific Countries - Germany & Canada

For some countries we have specific information about which stations to merge or split (and we hope that we obtain more of these lists as time goes on).  Currently we have information about German stations, whose id's start with 09 and 10 in the isd-history.txt files.  Here, the last 4 digits of the WMO-ID are important, and some stations have had their records split so that 09abcd and 10abcd are the same station.  So we use the selection algorithm to check these station-pairs specifically, and allow them to merge if they pass the same criteria as outlined above.

Including this information prior to the station selection criteria results in 8685 stations being selected compared to 8667 before.

For Canada, things are a little more complicated.  There are only 1000 WMO-IDs available for Canada, and as a result, stations with different locations have ended up with the same IDs.  Thanks to Environment Canada, we have a list of the station moves.  In this case we want to split up records so that apparent false mergers are not included in HadISD.  We are still working on including this information in the station selection code.

Note

These criteria and this procedure have not yet been finalised.  We may still revert to only merging/splitting stations where we have specific information.  If you have further suggestions or comments, please let us know.

Tuesday, 7 October 2014

Extending HadISD: Station Selection

I have started to re-assess the station selection part of HadISD.  During the early stages of creating HadISD (in around 2008), the ISD database was interrogated to find stations which would be suitable for HadISD.  This process, outlined in the paper, resulted in the 6103 stations which form HadISDv1.0.x.

However, this station list has not been updated since that point.  This means that we have not benefited from any new stations that have been added to the ISD database in recent years.  The static station list may also to be partly to blame for the jump in 2005 in the total number of stations with available data (see also HadISDH) and also the fall-off in the number of stations since 1990.
Fig. 1 - Number of stations which have data in any given year in HadISDv1.0.x.  These are all the individual input stations (including those merged to form composites), hence the peak is more than 6103.  The dip at 2005 is visible, as well as the drop before 1973 (which set the start period of HadISD v1.0.x).

At the same time as increasing the station selection we also intend to extend HadISD so that data is available and quality controlled prior to 1973.  It is clear from Fig. 1 why the start year of HadISDv1.0.x was chosen as 1973, however this does result in a relatively short period of record. 

Back to the beginning

So, we have gone back to the ISD database to dynamically return a station listing which could be run with each major update of HadISD.  Using the isd-history.txt we extracted those stations which have valid latitudes, longitudes and elevations, and also those which had at least 15 years between their start and end dates.  There are 29525 unique station IDs in the ISD database, and 14947 satisfy these criteria (these numbers will change fractionally as the ISD database is continually updated).

The isd-inventory.txt lists the number of observations in each month for each station.  We have used this to find those stations which report on average every 6 hours and which have observations in at least 15 years worth of months (180 months) to account for stations with many gaps.  This returns 8694 stations world wide.

Fig. 2. The number of stations which have data using the initial version of the updated station selection code.  The drops in 1972 and 2005 are still visible, but the gentle drop off from 1990 is less pronounced when compared to Fig. 1
We have dropped the reporting interval to every 6 hours rather than every 3 to try and select more stations in those regions where currently HadISD does not have many (e.g. central South America, Africa) but still maintain a subdaily resolution.

As can be seen in Fig. 2, there is still a large dip in 1972, and the drop in 2005 has also not entirely disappeared.  Some of these dips may be ameliorated by merging stations together.   However the drop off post 1990 is less prominent, and there are more stations overall.

To merge or not to merge?

As the ISD had many stations with short records, when creating HadISDv.1.0.x stations were merged to create ones with longer records.  This was done using a hierarchical table (see Table 1 in the paper) to identify potential candidates and then an in-depth and time-consuming manual process to reduce this to the mergers used.  If the station selection is to be run on each major update, then selecting these merger candidates would have to be automated.

We could go back to the raw ISD listings and find stations which are merging candidates with the ~8500 initially selected.  By merging these in, some of the gaps in Fig. 2 could be filled (but also possibly not).

As we have found over time, not all of these mergers are correct, and therefore a number of options present themselves:

  1. We can not merge at all, keep all the ISD station IDs as unique and be confident that by creating HadISD we have not degraded any of the data.
  2. We can merge only in cases where we have specific information (from a national met service, for example) as to the identity of stations.  This could also be applied in cases where we have information indicating that a split of a station record would be appropriate 
  3. We can merge (and split) when we have specific information, but also run an automated procedure to identify candidate stations to merge together.  An example algorithm has been produced by the International Surface Temperature Initiative databank v1.0 (see description in the paper).  However this approach is very likely to introduce spurious mergers, however careful we are with the algorithm.
We have not yet decided which route to follow, but are erring towards the second in the first instance.  

If you have any further suggestions or preferences, please leave a comment or get in touch.
 

Wednesday, 3 September 2014

Assessing the Homogeneity of HadISD

Apologies for the delay in writing this post, I have been distracted by working on the next version of our QC suite for the HadISD dataset.

Last month, our paper on the Pairwise Homogeneity Assessment of HadISD (Dunn et al, 2014) was published in Climate of the Past.  I have blogged about some of this work before here.

Pairwise Homogeneity Assessment

We used the Pairwise Homogenisation Algorithm (PHA) used for the US Historical Climate Network (USHCN) by Menne & Williams (2009).  This algorithm has the advantage of being able to run automatically for large networks of stations.  As there are ~6000 stations in HadISD, this was an important consideration when selecting which algorithm to use.  

The PHA has also been benchmarked for the USHCN stations (Williams et al, 2012) and also as part of an inter-algorithm comparison run as the COST-HOME project (Venema et al, 2012).  In the COST-HOME analysis, PHA was not the best performing, but was recommended as one of the algorithms to use when performing homogenisation on a monthly basis.  We would ideally like to be able to use a number of algorithms on the HadISD data to understand and be able to quantify (to some level) the uncertainties in the change point locations and adjustment magnitudes.  This will hopefully be available in future releases of HadISD.

As the paper is open access, I'm not going to go through the full details of the final methodology here.  However, if something is unclear in the paper, do leave a comment or get in touch.

Example Application - Global Land-Surface Temperatures


We decided from the outset that we would be unable to apply any of the adjustments found to the hourly data.  The adjustments were calculated using monthly averages, and so it is very likely that these are not going to be the correct values to use for the hourly data.  In fact, we may never be able to calculate appropriate adjustment magnitudes for this high time-resolution data as they are likely to depend on the cause of the inhomogeneity, the time of day and possibly even the weather type.  For example, the effect of moving a station will be different at different times of day (even with the same weather) as the sun hits the screen earlier or later; and if the weather changes, the new location may respond differently over the course of a day than the old.  Determining this from just the data itself (using no metadata as is the case for many stations) will be very hard.

So, what can be done with a list of change point dates and adjustment magnitudes.  Well, we thought that we could see what the "global" average land-surface temperature change from 1973-2013 is when taking stations with fewer and fewer or smaller and smaller inhomogeneities ("global" is in quotation marks because of the station distribution of HadISD - some regions are better sampled than others).   This would show (a) if there was an effect by excluding the more inhomogeneous stations, (b) what this effect was and (c) what other issues arise.

We were also inspired by the work of Callendar (1938, 1961) who used very few stations to estimate the global land-surface temperature and obtained results in agreement with the latest best estimates (see article by Ed Hawkins).  If Callendar's work obtained accurate results with few stations, then using a subsets of the HadISD stations should also.
Fig. 1. All 6103 stations from HadISD (black), CRUTEM4 (red) and CRUTEM4 with matched coverage (blue).  Bottom panel shows the deviations (with the full CRUTEM4 divided by 10 for clarity)

When using all 6103 stations there is already very good agreement between HadISD and the matched CRUTEM4 (Jones et al, 2012) - see Fig. 1.  The HadISD timeseries line (black) is not easily distinguished from the matched CRUTEM4 (blue), and the linear trends match very well.  We use linear trends in this case to easily summarise the change over the 40 years of data.  For this study, we use CRUTEM4 as our known comparison field.  Therefore we compare both against a version where we match the coverage, which is the best possible result we could achieve with HadISD, and the full version, which gives information as to what we are missing as a result of the incomplete coverage that HadISD has.
Fig. 2.  As for Fig. 1, but only taking stations where the largest inhomogeneity is less than 1C.
Only retaining stations where the largest inhomogeneity is <1C (4300 stations), improves the agreement, both for the linear trends but also for the deviations and root-mean-square errors (bottom panel).  Hence using the homogenisation assessment to select those stations which are relatively homogeneous does result in a better estimate.  However, the RMS against the full CRUTEM4 has increased from 0.13 to 0.17.
Fig.3.  As for Fig.1, but only taking stations where the largest inhomogeneity is less than 0.5C.
In Fig. 3 we were even more restrictive and here, although there was still a good match between HadISD and the matched CRUTEM4, the RMS has increased fractionally.  Similarly (not shown here, but in the paper) if we also restricted the number of change points that were detected in a station, the agreement also deteriorated, but not as rapidly as with the size of the inhomogeneity.
Fig. 4. As for Fig. 1, but only taking those stations in which no inhomogeneity was found or which could not be assessed.
Finally, in Fig. 4, we show the result when taking only those 1458 stations in which no inhomogeneity was found or which could not be assessed by PHA.  Here the agreement between HadISD and CRUTEM (both full and matched) is clearly deteriorating, although the linear trends do still agree within their uncertainties.  Also, the HadISD line does not emerge from the CRUTEM uncertainty envelope, even using these restrictions.


This indicates that eventually it is the coverage or which exact underlying stations are included which start to dominate any error, rather than the station quality itself.  However, removing the most inhomogeneous stations does result in better agreement, and also greater confidence on the part of researchers that results obtained are accurate.  Hence there is a balance to be drawn, and this will depend on the problem being addressed and also, at some level, to individual preference.  We can also use the results to double check our merging procedure as some inhomogeneities are likely to arise from stations that should not have been merged together.

References


Callendar, G. (1938). "The artificial production of carbon dioxide and its influence on temperature", Quarterly Journal of the Royal Meteorological Society, 64 (275), 223-240 DOI: 10.1002/qj.49706427503

Callendar, G. (1961). "Temperature fluctuations and trends over the earth", Quarterly Journal of the Royal Meteorological Society, 87 (371), 1-12 DOI: 10.1002/qj.49708737102

Dunn, R et al. (2014) "Pairwise homogeneity assessment of HadISD", Climate of the Past, 10, 1501

Hawkins, E & Jones, P (2013) "On increasing global temperatures: 75 years after Callendar", Quarterly Journal of the Royal Meteorological Society, 139, DOI: 10.1002/qj.2178

Jones, P et al. (2012) "Hemispheric and large-scale land-surface air temperature variations: An extensive revision and an update to 2010", Journal of Geophysical Research: Atmospheres", 117

Menne, M, & Williams, C, (2009), "Homogenization of temperature series via pairwise comparisons." Journal of Climate 22.7, 1700
 
Venema V et al, (2012) "Benchmarking homogenization algorithms for monthly data" Climate of the Past, 8,89
 
Williams C, N, Menne M, L and Thorne, P, W, (2012) "Benchmarking the performance of pairwise homogenization of surface temperatures in the United States" Journal of Geophysical Research, Atmospheres, 117

Monday, 9 June 2014

Cows, Milk and Heat Stress

We've just had a paper published in ERL on heat stress in UK dairy cattle and the effect this has on milk yields.  This was initially a short piece of work which at the time was the first application of HadISD to a specific project.  However the project grew and took longer than anticipated, and hence has only been published now.

(UK) dairy cattle and heat stress


Many animals are affected by high temperatures and humidities.  Think about a hot, humid day, personally I wouldn't have much energy and would prefer just lazing in the shade of a tree.  If, for example, I went for a run, I would get pretty warm, which may cause heat-stroke if I went for too long.  And if it is humid overnight, I don't sleep as well as normal either, which puts my body under more stress.  Cattle aren't often seen going for runs, but if they cannot cool themselves, they also get heat-stressed.  This means they divert energy from growing and producing milk to cooling down, and so you can measure the effect of the warm temperatures in the milk yields, which is handy as asking them how they feel is tricky!

We used a measure for heat stress called the "Temperature-Humidity Index" (THI) which combines the air temperature and the relative humidity.  Both of these are available from data in 68 of the HadISD stations used in this study.  When the THI rises above 70, then cattle experience heat-stress, and if it rises over ~90, then the heat stress is very severe and could be fatal.  We looked at the daily average THI, and across the UK, for most of the stations, this threshold of THI > 70 is only crossed in one or two days per year (Fig. 1).

Fig 1.  Top: The average number of days (over 1973-2012) with THI > 70 for the 68 UK stations.  Bottom: The number of days with THI > 70 in 2003.  Stations with no days where THI > 70 are shown with grey squares.  Note change in colour scale between the two panels.
The stations inside the M25 (the motorway around London) have more days, but we suspect that there are few cattle being raised that close to London.  Looking at the years of 2003 and 2006, which had very warm summers, there were 5-10 days where cattle could be heat-stressed (Fig. 1).  Using data provided by the Cattle Information Service we were able to look at data for milk yields on an animal-by-animal basis for a number of herds in the areas most likely to have been strongly affected by the high temperatures.

Factors affecting Milk Yields

There are a number of complicating factors when looking at milk yields from dairy cattle.  The amount of milk a cow produces depends on both the number of calves she has had, and also the number of days that have elapsed since her most recent calf was born (Fig 2.) 

Fig.2 Top: Yield vs days in milk, Bottom: Yield vs lactation number for herd 198 in Devon.  Blue points show individual measurements and the red ones show the average (over a 10 day bin or lactation number) with a 1-sigma uncertainty.


By looking at data on an animal-by-animal basis we were able to select cows in their first lactation and use ranges of 50 days since the birth of their calf to try and reduce the variation from these additional effects.  Fig. 3 shows the change in the milk yields for one herd over the entire record.  There is a clear decline in milk yields in 2006 for all ranges of days-in-milk.  The yield drops from 30 litres to between 15 and 20 litres.  There are few cattle contributing prior to 2004, resulting in a noisy curve and no clear indication of an effect in 2003.  Combined with the effect of heat-stress is also the reduction in pasture quality for grass fed cattle as the fields tend to dry out during hot weather.  However the feed information was not available

Fig. 3 Milk yields for one-calf cows in herd 4199 in Somerset.  The bands are the 1-sigma ranges for the three ranges of days-in-milk.  The dashed lines show the numbers of cows contributing to the mean milk yields.  The magenta line shows the Somerset average monthly temperature (Perry & Hollis, 2005).  The vertical dashed lines show the heatwaves of 2003 and 2006.

Climate Projections from UKCP09 

To study the effect of any future change in the climate on the number of days with high THI we used the projections from the UK Climate Prediction '09 assessment (UKCP09, Murphy et al 2009).  This is an 11-member ensemble of regional climate models runs on a 25 x 25 km grid.  The future climate is driven by a medium emissions scenario (A1B).  Fig. 4 shows the number of days in each grid box where the THI>70 for the south-west region.  Currently, as per the observations, there are only a few days on average where the threshold is exceeded.  But by the end of the century, this could have risen to around 30 days per year.
Fig. 4. The number of days per grid box where the THI > 70, averaged over the south-west region.  Each of the 11 ensemble members are shown in the grey lines, with the median and interquartile range being given by the thick black line and blue envelope respectively.

This could have a large effect on the milk yields produced by cows, and so impact on the viability of herds and dairy farming in parts of the UK.  Although keeping cattle indoors can mitigate the effect of direct solar radiation, the humidity in barns has been found to be always higher than indoors (Erbez et al, 2010).  Thought will have to be given in the future how best to keep cattle for their well being and also to ensure that dairy farming remains viable in parts of the UK.

References


 
Erbez M, Falta D and Chládek G 2010 The relationship between temperature and humidity outside and inside the permanently open-sided cows’ barn Acta Universitatis Agriculturae et Siliviculturae Mendelianae Brunensis (Brno, Česká Republika), LVIII 91–6
 
Murphy J M et al 2009 UK Climate Projections Science Report: Climate Change Projections Met Office Hadley Centre, Exeter, UK