Announcement

Collapse
No announcement yet.
X
  • Filter
  • Time
  • Show
Clear All
new posts

  • Winsorise

    Hello Sir/Madams,

    I have 6,071 firm-year observations for the period 2002-2014. I removed outlier firms from my sample by following the related literature.
    However, as I seen in the empirical literature, many people winsorise their sample at the 1st and 99th percentiles.
    When I try to winsorise my sample, I saw that my analysis results do not change, and any firm is not removed from the sample.
    Could you explain shortly, (1) what is the contribution of the winsorisation the sample that is done commonly by the empirical literature?
    (2) Which way is the more efficient: remove outliers from the sample or winsorisation the sample?

    Thanks for your concern and time.Regards,Hasan Tekin

  • #2
    There are some difficulties in answering this. Working backwards

    (2) What does "efficient" mean here? There is a strict statistical definition of efficiency; you may mean that or you may mean something else. I would assert that no discussion is possible without a clear definition.

    (1) "empirical literature": in what field? This is not an economics forum, and even if it were habits vary within economics, so far as I am aware, so you need to define your field.

    But Winsorizing at any percentile will remove nothing from the dataset as it just pulls in values more extreme to those percentiles.

    Comment


    • #3
      Hi Hasan,
      Winsorizing at the 1 and 99th percentile replaces values above 99th percentile and below the 1st percentile with the 99th and 1st percentile values.
      Usually winsorizing is done to remove outliers. Check your data thoroughly to understand if that is necessary.

      If you would like to drop outliers, you might want to use the trim option with winsor.

      Comment


      • #4
        Sam Basque:

        Some confusion here. winsor (SSC, my program) does not support a trim option and quite deliberately because I dislike the idea of removing outliers in this way and will not support it.

        You are presumably referring to some other program for which someone else is responsible.

        Comment


        • #5
          Hi Hasan,
          Nick is right.
          The trim option is for the winsor2 program. I suggested it since you wanted to remove firms from the sample.
          I, however, would not like to comment on methodological suitability.

          Comment


          • #6
            As usual, the best option is often to do all three and see if your results change considerably...

            Comment


            • #7
              Thanks for your answers and sorry for uninformative explanations.
              (1) My field is corporate finance. I analyse debt ratios and its determinants of UK firms.My sample has 9 firm-specific, 2 industry-specific, 3 macro-specific factors, and 3 dummy variables.
              (2) I mean here "could we say trimming outliers from the sample is better statistically than winsorising the sample or vice versa? Or Shall we do either trimming or winsorising the sample?

              Comment


              • #8
                (1) Thanks for the detail. The question is open to those working in that field, not me.

                (2) There is no universal agreement on this point. As Phil Bromiley often points out here, so many people do things like this that it's pretty much a standard method in fields that use it. My own line is that outliers are almost always best accommodated by working with some appropriate link (e.g. logarithm) but that view grows out of experience with environmental data, and may not travel well.

                Comment


                • #9
                  Let me add that my experience with economic data is similar to Nick's: the "outliers" are often informative about the functional form of the model and often it is possible to find a simple specification where what appeared to be outliers are not outliers at all.

                  To my mind, unless we have good reasons to suspect that there are errors in the observations, trimming or Winsorizing are just forms of making the data fit the model, which is the opposite of what we should be doing.

                  Joao

                  Comment


                  • #10
                    Thanks for the beneficial answers.
                    Regards.

                    Comment

                    Working...
                    X