Announcement

Collapse
No announcement yet.
X
  • Filter
  • Time
  • Show
Clear All
new posts

  • Converting Repeated Cross-Sectional Data into Panel Data

    Hi Everyone

    or one of my papers, I am using data from NFHS-4 and NFHS-5. Both surveys are cross-sectional, but they contain retrospective information with dates for several events. Using these retrospective dates, I converted each survey into a panel-type dataset and then merged the two reconstructed datasets into a single analytical dataset.
    The purpose is to answer a policy evaluation question using methods that are generally applied to panel or longitudinal data.
    I have attached two figures that summarise the methodology and data restructuring approach used in the paper.
    Click image for larger version

Name:	nfhs data transformation.png
Views:	6
Size:	340.7 KB
ID:	1787167 Click image for larger version

Name:	mcsp treatment timeline structure.png
Views:	4
Size:	282.6 KB
ID:	1787168



    I would be very grateful if you could advise me on the following points:
    • Is it statistically valid to restructure repeated cross-sectional survey data into a panel format when the underlying information is retrospective and time-stamped?
    • Can such a reconstructed dataset legitimately be treated as panel data for analysis?
    • More importantly, would it be methodologically appropriate to apply causal inference methods to such data, provided the assumptions of the chosen method are adequately addressed?
    • I am using the 'sdid' command for this panel data.
        My main concern is whether restructuring retrospective event histories into person-period observations provides a sufficiently valid basis for causal policy evaluation, or whether this should instead be considered a form of pseudo-panel or retrospective longitudinal analysis with important limitations.
    I would greatly value your opinion before proceeding further with the analysis and interpretation.
    Attached Files

  • #2
    Not a direct answer to your questions, but Fernando Rios-Avila - also active on this site - has a great github-page on the topics you are interested in (Fernando Rios-Avila – Playing With Stata). In particular so, the page "DID: Panel Data & Repeated Crossection" should be of relevance for your work: DID: Panel Data & Repeated Crossection – Playing With Stata

    Best, Frode

    Comment


    • #3
      Thank you for your response.

      Comment


      • #4
        As I see it: If recall is perfect, then it may be treated as typical panel data. The question is, therefore, whether recall is perfect? That's going to be a call, I believe, unless you have some reason to believe people might be engaging in expost rationalization and can test it.

        Comment


        • #5
          Pavan:
          potentially OoT but: is the sample measured on the same variables across time composed of the same units (set aside potential attrition-related issues)?
          Kind regards,
          Carlo
          (Stata 19.0)

          Comment


          • #6
            I agree with George that the main issue is whether there is lots of recall error. Probably there’s some, but it’s not clear it would systematically bias an intervention analysis. I do wonder if sdid is what you want. The fine print is that reliable inference requires many periods before and after the intervention date(s), and you wouldn’t have that. I’d recommend at least trying csdid or jwdid. If you think tends are different between treated and control — violation of parallel trends — you might try lwdid with the rolling(detrend) option.

            Comment


            • #7
              Originally posted by Carlo Lazzaro View Post
              Pavan:
              potentially OoT but: is the sample measured on the same variables across time composed of the same units (set aside potential attrition-related issues)?
              Hi Dr Carlo,

              The selected questions and responses are standarized and remained exactly the same in both the surveys.


              Thank You for your response.

              Comment


              • #8
                Originally posted by George Ford View Post
                As I see it: If recall is perfect, then it may be treated as typical panel data. The question is, therefore, whether recall is perfect? That's going to be a call, I believe, unless you have some reason to believe people might be engaging in expost rationalization and can test it.
                Hi Dr George,

                I agree that recall is an issue and we have addressed this issue in the limitation section of the paper. NFHS data collection procedure is pretty standarized and we are not too worried about the recall bias. For our analysis, we assume, that the recall was perfect and we are treating the data as it is.

                Given your response, I take it, technically our approach is correct and we can apply methods pertaining to panel data to the merged data.

                Anything else, I need to consider or address?


                Warm Regards
                Pavan.

                Comment


                • #9
                  Originally posted by Jeff Wooldridge View Post
                  I agree with George that the main issue is whether there is lots of recall error. Probably there’s some, but it’s not clear it would systematically bias an intervention analysis. I do wonder if sdid is what you want. The fine print is that reliable inference requires many periods before and after the intervention date(s), and you wouldn’t have that. I’d recommend at least trying csdid or jwdid. If you think tends are different between treated and control — violation of parallel trends — you might try lwdid with the rolling(detrend) option.


                  Dear Prof. Jeff,


                  Thank you for your response.


                  The unit for our analysis was month*state. We used monthly data for our analysis. This gave as roughly 60 units /periods before the intervention and 60 units/periods after the intervention. Thus, we decided to use the 'sdid'.

                  Once again, thank you for your inputs. I will use the command you suggested for additional analysis.

                  Given your inputs, I assume, that our approach is correct i.e., we can convert, event history given in two cross-sectional datasets into a single panel data and apply various types of difference-in-difference techniques on panel data?

                  Recall bias is an issue but we have addressed that in the limitations section of the paper and we are not too worried about it.

                  The most pressing issue for us is whether our approach is correct i.e., converting several repeated cross-section data into a single panel data and then using DiD techniques. We could not find any other paper that has adopted a similar approach in the field of public health, hence we are worried whether it will stand the scrutiny of methodological review.


                  Warm Regards

                  Pavan







                  Comment

                  Working...
                  X