如何基于csv中列中的多个描述符提取行计数，然后使用bash/python脚本导出新的csv？

Gene Element ---------- ---------- STBZIP1 G-box STBZIP1 G-box STBZIP1 MYC STBZIP1 MYC STBZIP1 MYC STBZIP10 MYC STBZIP10 MYC STBZIP10 MYC STBZIP10 G-box STBZIP10 G-box STBZIP10 G-box STBZIP10 G-box

2条回答

网友

1楼 · 编辑于 2024-10-02 00:42:03

因为您还要求提供bash版本，所以这里使用了awk¹。它是有注释的，而且输出的格式是“良好的”，所以代码有点大（大约20行没有注释）

awk '# First record line:
     # Storing all column names into elements, including
     # the first column name
     NR == 1 {firstcol=$1;element[$1]++}

     # Each line starting with the second one are datas
     # Occurrences are counted with an indexed array
     # count[x][y] contains the count of Element y for the Gene x
     NR > 2 {element[$2]++;count[$1][$2]++} 

     # Done, time for displaying the results
     END {
       # Let us display the first line, column names
       ## Left-justify the first col, because it is text
       printf "%-10s ", firstcol
       ## Other are counts, so we right-justify
       for (i in element) if (i != firstcol) printf "%10s ", i
       printf "\n"
       
       # Now an horizontal bar
       for (i in element) {
           c = 0
       while (c++ < 10) { printf "-"}
       printf " ";
       } 
       printf "\n"

       # Now, loop through the count records
       for (i in count) {
         # Left justification for the column name
         printf "%-10s ", i ;
         for(j in element)
           # For each counted element (ie except the first one),
           # print it right-justified
           if (j in count[i]) printf "%10s", count[i][j]
         printf "\n"
       }
     }' tab-separated-input.txt

结果:

Gene            G-box        MYC 
                  
STBZIP10            4         3
STBZIP1             2         3

¹由于Ed Morton

网友

2楼 · 编辑于 2024-10-02 00:42:03

文件格式为（此处名为input.csv）：

    Gene     Element   
             
  STBZIP1    G-box     
  STBZIP1    G-box     
  STBZIP1    MYC       
  STBZIP1    MYC       
  STBZIP1    MYC       
  STBZIP10   MYC       
  STBZIP10   MYC       
  STBZIP10   MYC       
  STBZIP10   G-box     
  STBZIP10   G-box     
  STBZIP10   G-box     
  STBZIP10   G-box

这个

import pandas as pd

df = pd.read_csv('input.csv', delim_whitespace=True, skiprows=1)
df.columns = ['Gene', 'Element']
df['Count'] = 1
df = df.pivot_table(index='Gene', columns='Element', aggfunc=sum)
print(df)

给你

         Count    
Element  G-box MYC
Gene              
STBZIP1      2   3
STBZIP10     4   3

相关问题更多 >

编程相关推荐

热门问题

热门文章

如何基于csv中列中的多个描述符提取行计数，然后使用bash/python脚本导出新的csv？

相关问题 更多 >

编程相关推荐

热门问题

热门文章

相关问题更多 >