Announcement

Collapse
No announcement yet.
X
  • Filter
  • Time
  • Show
Clear All
new posts

  • How to check if observation repeats in other columns

    Hello, I have a dataset with 4 edge lists next to one another.

    Code:
    * Example generated by -dataex-. For more info, type help dataex
    clear
    input long(start2018 end2018 start2019 end2019 start2020 end2020 start2021 end2021)
    2056 2637125 . . . . . .
    2056 1893381 . . . . . .
    2056 1970713 . . . . . .
    2056  533713 . . . . . .
    2056  885831 . . . . . .
    2056 1109190 . . . . . .
    2056  310034 . . . . . .
    2056 1987872 . . . . . .
    2056 1843845 . . . . . .
    2056 1863790 . . . . . .
    2056 2161689 . . . . . .
    2056 1675625 . . . . . .
    2056   29208 . . . . . .
    2056 1275296 . . . . . .
    2056 1893380 . . . . . .
    3213 2295372 . . . . . .
    3213 1448832 . . . . . .
    5909 2741756 . . . . . .
    5909 1532549 . . . . . .
    5909 2620465 . . . . . .
    end

    The columns start2019 and end2019 and other years take on values lower in the table, when all of the nodes of year 2018 are listed (there are about 10,000 words for each year). To perform a comparative social network analysis using nwcommands, I need networks with an equal amount of nodes. Thus, I want to omit nodes that do not appear in all four years.
    In other words, I need a way to identify and drop entries based on their appearance in every year's data. I want to ensure that each node (either starting or ending) appears in every year from 2018 to 2021 at least once (either in column 'startXXXX' or 'endXXXX'). If a node doesn't appear across all these years, it should be removed from the dataset entirely.

  • #2
    Have learned to do it in python instead as it turned out to be more straightforward.

    Comment

    Working...
    X